AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,131 results
Model Releases

Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks

DGX agent

arXiv:2604.18660v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in education, yet their default helpfulness often conflicts with pedagogical principles. Prior work

model-releasesarxiv-cs-ai
22 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Fast and Robust Diffusion Posterior Sampling for MR Image Reconstruction Using the Preconditioned Unadjusted Langevin Algorithm

DGX agent

arXiv:2512.05791v2 Announce Type: replace-cross Abstract: Purpose: The Unadjusted Langevin Algorithm (ULA) in combination with diffusion models can generate high quality MRI reconstructions with uncer

model-releasesarxiv-cs-cv
22 Apr 2026
Model Releases

FedProxy: Federated Fine-Tuning of LLMs via Proxy SLMs and Heterogeneity-Aware Fusion

DGX agent

arXiv:2604.19015v1 Announce Type: cross Abstract: Federated fine-tuning of Large Language Models (LLMs) is obstructed by a trilemma of challenges: protecting LLMs intellectual property (IP), ensuring

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Fine-tuning DeepSeek-OCR-2 for Molecular Structure Recognition

DGX agent

arXiv:2604.03476v2 Announce Type: replace-cross Abstract: Optical Chemical Structure Recognition (OCSR) is critical for converting 2D molecular diagrams from printed literature into machine-readable f

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Fine-Tuning Small Reasoning Models for Quantum Field Theory

DGX agent

arXiv:2604.18936v1 Announce Type: cross Abstract: Despite the growing application of Large Language Models (LLMs) to theoretical physics, there is little academic exploration into how domain-specific

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling

DGX agent

arXiv:2509.12052v3 Announce Type: replace Abstract: Current talking-head generation has gradually shifted from GAN-based methods to diffusion-based paradigms, achieving remarkable progress in visual f

model-releasesarxiv-cs-cv
22 Apr 2026
Model Releases

Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents

DGX agent

arXiv:2604.19457v1 Announce Type: new Abstract: Long-horizon enterprise agents make high-stakes decisions (loan underwriting, claims adjudication, clinical review, prior authorization) under lossy mem

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

From Experience to Skill: Multi-Agent Generative Engine Optimization via Reusable Strategy Learning

DGX agent

arXiv:2604.19516v1 Announce Type: new Abstract: Generative engines (GEs) are reshaping information access by replacing ranked links with citation-grounded answers, yet current Generative Engine Optimi

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

From Natural Language to Executable Narsese: A Neuro-Symbolic Benchmark and Pipeline for Reasoning with NARS

DGX agent

arXiv:2604.18873v1 Announce Type: new Abstract: Large language models (LLMs) are highly capable at language generation, but they remain unreliable when reasoning requires explicit symbolic structure,

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models

DGX agent

arXiv:2511.10899v2 Announce Type: replace Abstract: Tool-augmented Language Models (TaLMs) can invoke external tools to solve problems beyond their parametric capacity. However, it remains unclear whe

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

Gated Memory Policy

DGX agent

arXiv:2604.18933v1 Announce Type: cross Abstract: Robotic manipulation tasks exhibit varying memory requirements, ranging from Markovian tasks that require no memory to non-Markovian tasks that depend

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

GenerativeMPC: VLM-RAG-guided Whole-Body MPC with Virtual Impedance for Bimanual Mobile Manipulation

DGX agent

arXiv:2604.19522v1 Announce Type: new Abstract: Bimanual mobile manipulation requires a seamless integration between high-level semantic reasoning and safe, compliant physical interaction - a challeng

model-releasesarxiv-cs-ro
22 Apr 2026
Model Releases

GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines

DGX agent

arXiv:2508.06226v2 Announce Type: replace Abstract: Geometry problem solving (GPS) poses significant challenges for Multimodal Large Language Models (MLLMs) in diagram comprehension, knowledge applica

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

GRAFT: Geometric Refinement and Fitting Transformer for Human Scene Reconstruction

DGX agent

arXiv:2604.19624v1 Announce Type: new Abstract: Reconstructing physically plausible 3D human-scene interactions (HSI) from a single image currently presents a trade-off: optimization based methods off

model-releasesarxiv-cs-cv
22 Apr 2026
Model Releases

GRASPrune: Global Gating for Budgeted Structured Pruning of Large Language Models

DGX agent

arXiv:2604.19398v1 Announce Type: new Abstract: Large language models (LLMs) are expensive to serve because model parameters, attention computation, and KV caches impose substantial memory and latency

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models

DGX agent

arXiv:2604.19300v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks. However, hallucination, where models

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human-LLM Collaborative Writing

DGX agent

arXiv:2604.19274v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as co-authors in collaborative writing, where users begin with rough drafts and rely on LLMs to compl

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams

DGX agent

arXiv:2604.18901v1 Announce Type: cross Abstract: Harmful intent is geometrically recoverable from large language model residual streams: as a linear direction in most layers, and as angular deviation

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

HarmoniDiff-RS: Training-Free Diffusion Harmonization for Satellite Image Composition

DGX agent

arXiv:2604.19392v1 Announce Type: new Abstract: Satellite image composition plays a critical role in remote sensing applications such as data augmentation, disaste simulation, and urban planning. We p

model-releasesarxiv-cs-cv
22 Apr 2026
Model Releases

Has Automated Essay Scoring Reached Sufficient Accuracy? Deriving Achievable QWK Ceilings from Classical Test Theory

DGX agent

arXiv:2604.19131v1 Announce Type: new Abstract: Automated essay scoring (AES) is commonly evaluated on public benchmarks using quadratic weighted kappa (QWK). However, because benchmark labels are ass

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

HELM: Harness-Enhanced Long-horizon Memory for Vision-Language-Action Manipulation

DGX agent

arXiv:2604.18791v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models fail systematically on long-horizon manipulation tasks despite strong short-horizon performance. We show that this

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Heterogeneity-Aware Personalized Federated Learning for Industrial Predictive Analytics

DGX agent

arXiv:2604.19451v1 Announce Type: new Abstract: Federated prognostics enable clients (e.g., companies, factories, and production lines) to collaboratively develop a failure time prediction model while

model-releasesarxiv-cs-lg
22 Apr 2026
Model Releases

How Adversarial Environments Mislead Agentic AI?

DGX agent

arXiv:2604.18874v1 Announce Type: new Abstract: Tool-integrated agents are deployed on the premise that external tools ground their outputs in reality. Yet this very reliance creates a critical attack

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

How Far Are Video Models from True Multimodal Reasoning?

DGX agent

arXiv:2604.19193v1 Announce Type: new Abstract: Despite remarkable progress toward general-purpose video models, a critical question remains unanswered: how far are these models from achieving true mu

model-releasesarxiv-cs-cv
22 Apr 2026
Model Releases

HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing

DGX agent

arXiv:2604.19071v1 Announce Type: new Abstract: Evaluating the writing capabilities of large language models (LLMs) remains a significant challenge due to the multidimensional nature of writing skills

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

HP-Edit: A Human-Preference Post-Training Framework for Image Editing

DGX agent

arXiv:2604.19406v1 Announce Type: cross Abstract: Common image editing tasks typically adopt powerful generative diffusion models as the leading paradigm for real-world content editing. Meanwhile, alt

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Human-Guided Harm Recovery for Computer Use Agents

DGX agent

arXiv:2604.18847v1 Announce Type: new Abstract: As LM agents gain the ability to execute actions on real computer systems, we need ways to not only prevent harmful actions at scale but also effectivel

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Improving the Distributional Alignment of LLMs using Supervision

DGX agent

arXiv:2507.00439v4 Announce Type: replace Abstract: The ability to accurately align LLMs with diverse population groups on subjective questions would have great value. In this work, we show that addin

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text

DGX agent

arXiv:2604.19298v1 Announce Type: cross Abstract: We introduce IndiaFinBench, to our knowledge the first publicly available evaluation benchmark for assessing large language model (LLM) performance on

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

InsideOut: Measuring and Mitigating Insider-Outsider Bias in Interview Script Generation

DGX agent

arXiv:2509.21080v2 Announce Type: replace-cross Abstract: Advancements in Large language models (LLMs) have enabled a variety of downstream applications like story and interview script generation. How

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Intentional Updates for Streaming Reinforcement Learning

DGX agent

arXiv:2604.19033v1 Announce Type: cross Abstract: In gradient-based learning, a step size chosen in parameter units does not produce a predictable per-step change in function output. This often leads

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

LatticeVision: Image to Image Networks for Modeling Non-Stationary Spatial Data

DGX agent

arXiv:2505.09803v3 Announce Type: replace-cross Abstract: In many applications, we wish to fit a parametric statistical model to a small ensemble of spatially distributed random variables ('fields').

model-releasesarxiv-cs-lg
22 Apr 2026
Model Releases

LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification

DGX agent

arXiv:2604.18878v1 Announce Type: new Abstract: We introduce LegalBench-BR, the first public benchmark for evaluating language models on Brazilian legal text classification. The dataset comprises 3,10

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

LePREC: Reasoning as Classification over Structured Factors for Assessing Relevance of Legal Issues

DGX agent

arXiv:2604.19464v1 Announce Type: cross Abstract: More than half of the global population struggles to meet their civil justice needs due to limited legal resources. While Large Language Models (LLMs)

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Less Is More: Cognitive Load and the Single-Prompt Ceiling in LLM Mathematical Reasoning

DGX agent

arXiv:2604.18897v1 Announce Type: new Abstract: We present a systematic empirical study of prompt engineering for formal mathematical reasoning in the context of the SAIR Equational Theories Stage 1 c

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models

DGX agent

arXiv:2601.22737v2 Announce Type: replace Abstract: The robust safety of Vision-Language Large Models (VLLMs) against joint multilingual and multimodal threats remains severely underexplored. Current

model-releasesarxiv-cs-cv
22 Apr 2026
Model Releases

LiveVLN: Breaking the Stop-and-Go Loop in Vision-Language Navigation

DGX agent

arXiv:2604.19536v1 Announce Type: new Abstract: Recent navigation systems achieve strong benchmark results, yet real-world deployment often remains visibly stop-and-go. This bottleneck arises because

model-releasesarxiv-cs-ro
22 Apr 2026
Model Releases

LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models

DGX agent

arXiv:2604.18803v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly deployed in settings where reliable visual grounding carries operational consequences, yet their behavi

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Lost in Translation: Do LVLM Judges Generalize Across Languages?

DGX agent

arXiv:2604.19405v1 Announce Type: new Abstract: Automatic evaluators such as reward models play a central role in the alignment and evaluation of large vision-language models (LVLMs). Despite their gr

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

LoViF 2026 Challenge on Real-World All-in-One Image Restoration: Methods and Results

DGX agent

arXiv:2604.19445v1 Announce Type: new Abstract: This paper presents a review for the LoViF Challenge on Real-World All-in-One Image Restoration. The challenge aimed to advance research on real-world a

model-releasesarxiv-cs-cv
22 Apr 2026
Model Releases

LSTM-MAS: A Long Short-Term Memory Inspired Multi-Agent System for Long-Context Understanding

DGX agent

arXiv:2601.11913v2 Announce Type: replace-cross Abstract: Effectively processing long contexts remains a fundamental yet unsolved challenge for large language models (LLMs). Existing single-LLM-based

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Mango: Multi-Agent Web Navigation via Global-View Optimization

DGX agent

arXiv:2604.18779v1 Announce Type: new Abstract: Existing web agents typically initiate exploration from the root URL, which is inefficient for complex websites with deep hierarchical structures. Witho

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

Mechanistic Anomaly Detection via Functional Attribution

DGX agent

arXiv:2604.18970v1 Announce Type: new Abstract: We can often verify the correctness of neural network outputs using ground truth labels, but we cannot reliably determine whether the output was produce

model-releasesarxiv-cs-lg
22 Apr 2026
Model Releases

Micro Language Models Enable Instant Responses

DGX agent

arXiv:2604.19642v1 Announce Type: new Abstract: Edge devices such as smartwatches and smart glasses cannot continuously run even the smallest 100M-1B parameter language models due to power and compute

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

DGX agent

arXiv:2511.11793v3 Announce Type: replace Abstract: We present MiroThinker v1.0, an open-source research agent designed to advance tool-augmented reasoning and information-seeking capabilities. Unlike

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation

DGX agent

arXiv:2604.18914v1 Announce Type: cross Abstract: While multilingual large language models (LLMs) perform well on high-level tasks like translation and question answering, their ability to handle gram

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

MSDS: Deep Structural Similarity with Multiscale Representation

DGX agent

arXiv:2604.19159v1 Announce Type: new Abstract: Deep-feature-based perceptual similarity models have demonstrated strong alignment with human visual perception in Image Quality Assessment (IQA). Howev

model-releasesarxiv-cs-cv
22 Apr 2026
Model Releases

Multi-Domain Learning with Global Expert Mapping

DGX agent

arXiv:2604.18842v1 Announce Type: new Abstract: Human perception generalizes well across different domains, but most vision models struggle beyond their training data. This gap motivates multi-dataset

model-releasesarxiv-cs-cv
22 Apr 2026
← Previous
1…309310311312313…357
Next →