AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

Content type
86,510Total entries
1Added by human
86,509Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
50,764 results
Model Releases

Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training

DGX agent

arXiv:2605.09608v1 Announce Type: new Abstract: Continual post-training aims to extend large language models (LLMs) with new knowledge, skills, and behaviors, yet it remains unclear when sequential up

model-releasesarxiv-cs-lg
12 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal

DGX agent

arXiv:2605.09502v1 Announce Type: cross Abstract: Chain-of-thought (CoT) prompting assumes that generated reasoning reflects a model's internal computation. We show this assumption is wrong in a speci

model-releasesarxiv-cs-ai
12 May 2026
Research

Hypothesis-Driven Deep Research with Large Language Models: A Structured Methodology for Automated Knowledge Discovery

DGX agent

arXiv:2605.10224v1 Announce Type: new Abstract: Current AI-powered research systems adopt a direct search-then-summarize paradigm that treats hypotheses as end products of scientific discovery. We arg

researcharxiv-cs-ai
12 May 2026
Research

Infinite Mask Diffusion for Few-Step Distillation

DGX agent

arXiv:2605.10518v1 Announce Type: cross Abstract: Masked Diffusion Models (MDMs) have emerged as a promising alternative to autoregressive models in language modeling, offering the advantages of paral

researcharxiv-cs-ai
12 May 2026
Model Releases

Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation

DGX agent

arXiv:2601.11258v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) face the 'knowledge cutoff' challenge, where their frozen parametric memory prevents direct internalization of ne

model-releasesarxiv-cs-ai
12 May 2026
Research

Learning the Interaction Prior for Protein-Protein Interaction Prediction: A Model-Agnostic Approach

DGX agent

arXiv:2605.09964v1 Announce Type: new Abstract: Protein-protein interactions (PPIs) are fundamental to cellular function and disease mechanisms. Current learning-based PPI predictors focus on learning

researcharxiv-cs-ai
12 May 2026
Model Releases

MIDUS: Memory-Infused Depth Up-Scaling

DGX agent

arXiv:2512.13751v2 Announce Type: replace-cross Abstract: Expanding pre-trained language models offers a practical way to increase capacity without training larger models from scratch. Depth Up-Scalin

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims

DGX agent

arXiv:2605.08192v1 Announce Type: cross Abstract: Frontier AI safety claims - published assertions that a highly capable general-purpose model is below a threshold of concern, adequately mitigated, or

model-releasesarxiv-cs-ai
12 May 2026
Research

Nix and Fix: Targeting 1000x Compression of 3D Gaussian Splatting with Diffusion Models

DGX agent

arXiv:2602.04549v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) revolutionized novel view rendering. Instead of inferring from dense spatial points, as implicit representations do, 3D

researcharxiv-cs-cv
12 May 2026
Model Releases

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness

DGX agent

arXiv:2605.10379v1 Announce Type: new Abstract: Large language models (LLMs) have become capable mathematical problem-solvers, often producing correct proofs for challenging problems. However, correct

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

DGX agent

arXiv:2605.08762v1 Announce Type: cross Abstract: Current omni-modal benchmarks mainly evaluate models under settings where multiple modalities are provided simultaneously, while the ability to start

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Overconfident and Blind to Details: Fixing Prompt Insensitivity with Abductive Preference Learning

DGX agent

arXiv:2510.09887v2 Announce Type: replace Abstract: Vision and language models frequently ignore semantically critical input edits, defaulting to pretraining priors. For example, models will confident

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Pix2Fact: When Vision Is Not Enough -- Benchmarking Fine-Grained VQA with Web Verification on High-Resolution Real-World Scenes

DGX agent

arXiv:2602.00593v2 Announce Type: replace Abstract: Despite progress on general tasks, vision-language models (VLMs) still struggle with challenges that demand both fine-grained visual grounding and e

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark

DGX agent

arXiv:2410.14702v2 Announce Type: replace Abstract: Multi-modal Large Language Models (MLLMs) exhibit impressive problem-solving abilities in various domains, but their visual comprehension and abstra

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

ReorgGS: Equivalent Distribution Reorganization for 3D Gaussian Splatting

DGX agent

arXiv:2605.08739v1 Announce Type: new Abstract: A converged 3D Gaussian Splatting (3DGS) model may approximate the target scene while remaining poorly parameterized for further optimization. We identi

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs

DGX agent

arXiv:2605.09239v1 Announce Type: new Abstract: Large language models fail at counting repeated tokens despite strong performance on broader reasoning benchmarks. These failures are commonly attribute

model-releasesarxiv-cs-cl
12 May 2026
Safety

Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks

DGX agent

arXiv:2605.08257v1 Announce Type: cross Abstract: Motivated by the challenge to improve the adversarial robustness, security, and trust of medical decision making intelligent agents, this study develo

safetyarxiv-cs-ai
12 May 2026
Model Releases

Seed Hijacking of LLM Sampling and Quantum Random Number Defense

DGX agent

arXiv:2605.08313v1 Announce Type: cross Abstract: Large language models (LLMs) rely on deterministic pseudorandom number generators (PRNGs) for autoregressive sampling, creating a critical supply-chai

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

SEMASIA: A Large-Scale Dataset of Semantically Structured Latent Representations

DGX agent

arXiv:2605.09485v1 Announce Type: new Abstract: Latent representations learned by neural networks often exhibit semantic structure, where concept similarity is reflected by geometric proximity in embe

model-releasesarxiv-cs-lg
12 May 2026
Research

Sequential Membership Inference Attacks

DGX agent

arXiv:2602.16596v2 Announce Type: replace Abstract: Modern AI models are not static. They go through multiple updates in their lifecycles. We propose to design Sequential Membership Inference (SeMI) a

researcharxiv-cs-lg
12 May 2026
Model Releases

SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy

DGX agent

arXiv:2605.09598v1 Announce Type: new Abstract: Vision-language models (VLMs) have recently shown strong potential in soccer video understanding. However, given the high complexity of soccer videos du

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

DGX agent

arXiv:2605.09063v1 Announce Type: new Abstract: Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challengi

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

SPDEBench: An Extensive Benchmark for Learning Stochastic PDEs

DGX agent

arXiv:2505.18511v2 Announce Type: replace Abstract: Stochastic Partial Differential Equations (SPDEs) driven by random noise play a central role in modeling physical processes with rough spatio-tempor

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Statistical Scouting Finds Debate-Safe but Not Debate-Useful Cases: A Matched-Ceiling Study of Open-Weight LLM Reasoning Protocols

DGX agent

arXiv:2605.09618v1 Announce Type: new Abstract: When should a language model answer directly, sample and vote, or engage in multi-agent debate? Recent work shows voting often explains much of the gain

model-releasesarxiv-cs-cl
12 May 2026
Research

Test-Time Training for Visual Foresight Vision-Language-Action Models

DGX agent

arXiv:2605.08215v1 Announce Type: new Abstract: Visual Foresight VLA (VF-VLA) has become a prominent architectural choice in the recent VLA due to its impressive performance. Nevertheless, the inheren

researcharxiv-cs-cv
12 May 2026
Model Releases

The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark

DGX agent

arXiv:2605.09900v1 Announce Type: new Abstract: A vision-language model can look at a knot diagram and report what it sees, yet fail to act on that structure. KnotBench pairs an 858,318-image corpus f

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why

DGX agent

arXiv:2605.10889v1 Announce Type: cross Abstract: On-policy distillation offers dense, per-token supervision for training reasoning models; however, it remains unclear under which conditions this sign

model-releasesarxiv-cs-ai
12 May 2026
Research

Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation

DGX agent

arXiv:2605.09981v1 Announce Type: cross Abstract: Multimodal models that jointly reason over protein sequences, structures, and function annotations within a unified representation hold immense potent

researcharxiv-cs-ai
12 May 2026
Model Releases

Bayesian Fine-tuning in Projected Subspaces

DGX agent

arXiv:2605.07706v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning of large models by decomposing weight updates into low-rank matrices, significantly r

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight

DGX agent

arXiv:2605.07021v1 Announce Type: new Abstract: Reasoning in Large Language Models (LLMs) poses a challenge for oversight as many misaligned behaviors do not surface until reasoning concludes. To addr

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

BoHA: Blockwise Hadamard Product Adaptation for Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2509.21637v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning (PEFT) of large language models trains a small task-specific parameter set while keeping the pretrained model frozen

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Can You Break RLVER? Probing Adversarial Robustness of RL-Trained Empathetic Agents

DGX agent

arXiv:2605.07138v1 Announce Type: new Abstract: Reinforcement learning from verifiable emotion rewards RLVER has produced language models with strong empathetic performance, evaluated on benchmarks th

model-releasesarxiv-cs-ai
11 May 2026
Hardware

Closed-Form Linear-Probe Dataset Distillation for Pre-trained Vision Models

DGX agent

arXiv:2605.07194v1 Announce Type: cross Abstract: Dataset distillation compresses a large training set into a small synthetic set that preserves downstream training utility. While most existing method

hardwarearxiv-cs-ai
11 May 2026
Model Releases

Delulu: A Verified Multi-Lingual Benchmark for Code Hallucination Detection in Fill-in-the-Middle Tasks

DGX agent

arXiv:2605.07024v1 Announce Type: new Abstract: Large Language Models for code generation frequently produce hallucinations in Fill-in-the-Middle (FIM) tasks -- plausible but incorrect completions suc

model-releasesarxiv-cs-lg
11 May 2026
Safety

Dynamic one-time delivery of critical data by small and sparse UAV swarms: a model problem for MARL scaling studies

DGX agent

arXiv:2512.09682v2 Announce Type: replace-cross Abstract: This work studies the application of Multi-Agent Reinforcement Learning (MARL) to decentralized control of unmanned aerial vehicles to relay a

safetyarxiv-cs-ai
11 May 2026
Model Releases

FastOmniTMAE: Parallel Clause Learning for Scalable and Hardware-Efficient Tsetlin Embeddings

DGX agent

arXiv:2605.06982v1 Announce Type: new Abstract: Embedding models in natural language processing (NLP) increasingly rely on deep architectures such as BERT, while simpler models such as Word2Vec provid

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

GLiGuard: Schema-Conditioned Classification for LLM Safeguard

DGX agent

arXiv:2605.07982v1 Announce Type: new Abstract: Ensuring safe, policy-compliant outputs from large language models requires real-time content moderation that can scale across multiple safety dimension

model-releasesarxiv-cs-cl
11 May 2026
Model Releases

Inference Time Causal Probing in LLMs

DGX agent

arXiv:2605.07631v1 Announce Type: new Abstract: Causal probing methods aim to test and control how internal representations influence the behavior of generative models. In causal probing, an intervent

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

IntentGrasp: A Comprehensive Benchmark for Intent Understanding

DGX agent

arXiv:2605.06832v1 Announce Type: cross Abstract: Accurately understanding the intent behind speech, conversation, and writing is crucial to the development of helpful Large Language Model (LLM) assis

model-releasesarxiv-cs-ai
11 May 2026
Applications

Large Video Planner Enables Generalizable Robot Control

DGX agent

arXiv:2512.15840v2 Announce Type: replace-cross Abstract: General-purpose robots require decision-making models that generalize across diverse tasks and environments. Recent works build robot foundati

applicationsarxiv-cs-cv
11 May 2026
Model Releases

MathlibPR: Pull Request Merge-Readiness Benchmark for Formal Mathematical Libraries

DGX agent

arXiv:2605.07147v1 Announce Type: cross Abstract: The ecosystem of Lean and Mathlib has become the de facto standard for large language model (LLM) assisted formal reasoning with remarkable successes

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

MAVEN: Multi-Agent Verification-Elaboration Network with In-Step Epistemic Auditing

DGX agent

arXiv:2605.07646v1 Announce Type: cross Abstract: While explicit reasoning trajectories enhance model interpretability, existing paradigms often rely on monolithic chains that lack intermediate verifi

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Neural Neural Scaling Laws

DGX agent

arXiv:2601.19831v2 Announce Type: replace-cross Abstract: Neural scaling laws predict how language model performance improves with increased training inputs. While aggregate metrics like validation lo

model-releasesarxiv-cs-cl
11 May 2026
Model Releases

Rep2Text: Decoding Full Text from a Single LLM Token Representation

DGX agent

arXiv:2511.06571v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have achieved remarkable progress across diverse tasks, yet their internal mechanisms remain largely opaque. In t

model-releasesarxiv-cs-ai
11 May 2026
Research

Rethinking Experience Utilization in Self-Evolving Language Model Agents

DGX agent

arXiv:2605.07164v1 Announce Type: new Abstract: Self-evolving agents improve by accumulating and reusing experience from past interactions. Existing work has largely focused on how experience is const

researcharxiv-cs-cl
11 May 2026
Model Releases

SR^2-LoRA: Self-Rectifying Inter-layer Relations in Low-Rank Adaptation for Class-Incremental Learning

DGX agent

arXiv:2605.07420v1 Announce Type: cross Abstract: Pre-trained models with parameter-efficient fine-tuning (PEFT) have demonstrated promising potential for class-incremental learning (CIL), yet catastr

model-releasesarxiv-cs-cv
11 May 2026
Model Releases

Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts

DGX agent

arXiv:2605.07395v1 Announce Type: cross Abstract: Efficient routing across multiple LLMs enables cost-quality tradeoffs by directing queries to the cheapest capable model. Prior work attributes routin

model-releasesarxiv-cs-ai
11 May 2026
Agents

Agentic publications: redesigning scientific publishing in the age of thinking large language models

DGX agent

arXiv:2505.13246v2 Announce Type: replace Abstract: Purpose: This paper introduces the concept of 'Agentic Publication,' a novel LLM-driven framework designed to complement traditional scientific publ

agentsarxiv-cs-ai
7 May 2026
← Previous
1…291292293294295…1058
Next →