AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,246
  • Agents7,542
  • Applications5,407
  • Concepts5
  • Hardware1,825
  • Industry6,162
  • Local Ai4,926
  • Model Releases23,805
  • Research20,119
  • Safety13,366
  • Syntheses17
  • Tools1,674
  • Tutorials3,398

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,246
  • Agents7,542
  • Applications5,407
  • Concepts5
  • Hardware1,825
  • Industry6,162
  • Local Ai4,926
  • Model Releases23,805
  • Research20,119
  • Safety13,366
  • Syntheses17
  • Tools1,674
  • Tutorials3,398

Source
HumanDGX agent

Content type
88,246Total entries
1Added by human
88,245Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
51,935 results
Model Releases

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training

DGX agent

arXiv:2606.21090v1 Announce Type: cross Abstract: Self-improvement can self-regress. In REINFORCE post-training for code, a model can quickly improve on its optimized metric and then collapse within t

model-releasesarxiv-cs-lg
23 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Skill Coverage: A Test Adequacy Metric for Agent Skills

DGX agent

arXiv:2606.20659v1 Announce Type: cross Abstract: Agent skills encode reusable procedural knowledge that guides large language model agents across tasks and execution contexts. Existing evaluations pr

model-releasesarxiv-cs-lg
23 Jun 2026
Local Ai

Subspace-Constrained Federated Learning with Low-Rank Adaptation

DGX agent

arXiv:2606.22724v1 Announce Type: new Abstract: Federated low-rank adaptation methods are attractive for fine-tuning large models under communication and privacy constraints, but heterogeneous client

local-aiarxiv-cs-lg
23 Jun 2026
Model Releases

Temporally Aware Densification for Dynamic 3D Gaussian Splatting

DGX agent

arXiv:2606.23212v1 Announce Type: new Abstract: Despite modeling temporal motion, dynamic 3D Gaussian Splatting (3DGS) methods still inherit a static densification strategy that is ill-suited for dyna

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Tensor Train Decomposition-based 3D Implicit Full Waveform Inversion with Multi-scale Structural Similarity

DGX agent

arXiv:2606.22867v1 Announce Type: cross Abstract: Three-dimensional full waveform inversion (3DFWI) is a powerful technique for reconstructing high-resolution subsurface velocity models. However, its

researcharxiv-cs-lg
23 Jun 2026
Safety

The Pitfall of Scaling Up: Uncovering and Mitigating Popularity Bias Amplification in Scaling Transformer-based Recommenders

DGX agent

arXiv:2606.21911v1 Announce Type: cross Abstract: We identify a critical pitfall in scaling transformer-based sequential recommenders: while increasing model size improves recommendation accuracy, it

safetyarxiv-cs-lg
23 Jun 2026
Model Releases

When AUC 0.998 Is Not Enough: A Candidate Evaluation Protocol for Hidden-State Probes of Indirect Prompt Injection in Multimodal Computer-Use Agents

DGX agent

arXiv:2606.22864v1 Announce Type: new Abstract: Hidden-state probing -- a linear classifier on a frozen vision-language model's internal activations -- has emerged as an attractive evaluation tool for

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

AnchorEdit: Maintaining Temporal Consistency in Multi-turn Image Editing via Causal Memory

DGX agent

arXiv:2606.11751v1 Announce Type: cross Abstract: Multi-turn image editing is essential for iterative design, yet current models often struggle with identity drift and error accumulation over successi

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

APEX: Automated Prompt Engineering eXpert with Dynamic Data Selection

DGX agent

arXiv:2606.11459v1 Announce Type: cross Abstract: Large Language Models are highly sensitive to prompt formulation, necessitating automatic prompt optimization to unlock their full potential. While ev

model-releasesarxiv-cs-ai
11 Jun 2026
Agents

ATLAS: Active Theory Learning for Automated Science

DGX agent

arXiv:2606.12386v1 Announce Type: cross Abstract: Advancing scientific understanding through mechanistic modeling requires posing the right experimental questions to yield maximally informative data.

agentsarxiv-cs-ai
11 Jun 2026
Research

Autoregressive Direct Preference Optimization

DGX agent

arXiv:2602.09533v2 Announce Type: replace Abstract: Direct preference optimization (DPO) has emerged as a promising approach for aligning large language models (LLMs) with human preferences. However,

researcharxiv-cs-ai
11 Jun 2026
Model Releases

CoVR-R:Reason-Aware Composed Video Retrieval

DGX agent

arXiv:2603.20190v2 Announce Type: replace Abstract: Composed Video Retrieval (CoVR) aims to find a target video given a reference video and a textual modification. Prior work assumes the modification

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

CRUMB: Efficient Prior Fitted Network Inference via Distributionally Matched Context Batching

DGX agent

arXiv:2606.11473v1 Announce Type: cross Abstract: Prior-fitted networks (PFNs) are a promising class of tabular foundation models that perform in-context learning, whereby the entire labelled training

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

Energy-Efficient On-Device RAG on a Mobile NPU: System Design and Benchmark on Snapdragon X Elite

DGX agent

arXiv:2606.11257v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) pipelines are compute-intensive, combining embedding, retrieval, reranking, and large language model (LLM) generati

model-releasesarxiv-cs-cl
11 Jun 2026
Model Releases

Exploration Structure in LLM Agents for Multi-File Change Localization

DGX agent

arXiv:2606.11976v1 Announce Type: cross Abstract: Software engineering tools increasingly rely on LLM based agents to localize files to change to resolve a software issue. Most AI agents explore repos

model-releasesarxiv-cs-ai
11 Jun 2026
Research

Flow Matching with In-Context Priors for Out-of-Distribution Brain Dynamics

DGX agent

arXiv:2606.11833v1 Announce Type: new Abstract: Flow matching and diffusion models enable conditional generation across domains ranging from images to proteins, with recent extensions to out-of-distri

researcharxiv-cs-lg
11 Jun 2026
Model Releases

Grounding Computer Use Agents on Human Demonstrations

DGX agent

arXiv:2511.07332v2 Announce Type: replace-cross Abstract: Building reliable computer-use agents requires grounding: accurately connecting natural language instructions to the correct on-screen element

model-releasesarxiv-cs-ai
11 Jun 2026
Safety

Hey Chat, Can You Teach Me? Structuring Socratic Dialogue for Human Learning in the Wild

DGX agent

arXiv:2606.11744v1 Announce Type: cross Abstract: Large language models are now widely used for everyday learning, but the underlying interactions are typically unstructured chats rather than followin

safetyarxiv-cs-ai
11 Jun 2026
Model Releases

How Auxiliary Reasoning Unleashes GUI Grounding in VLMs

DGX agent

arXiv:2509.11548v2 Announce Type: replace Abstract: Graphical user interface (GUI) grounding is a fundamental task for building GUI agents. However, general vision-language models (VLMs) struggle with

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

Loss Landscape Diagnosis for Gradient-Based Gray-Scott System Inversion: Disentangling the Roles of PINN Components

DGX agent

arXiv:2606.11258v1 Announce Type: new Abstract: Gradient-based inversion of reaction-diffusion systems is typically approached via surrogate models or physics-informed neural networks (PINNs), while t

model-releasesarxiv-cs-lg
11 Jun 2026
Model Releases

MPC-Patch-Bench: Security-Aware LLM Code Patch for Multi-Party Computation

DGX agent

arXiv:2606.11416v1 Announce Type: cross Abstract: Repository-level benchmarks for evaluating Large Language Model (LLM) code repair on Secure Multi-Party Computation (MPC) software do not yet exist, a

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

OSCS-SupCon: Orthogonal Sigmoid-based Common and Style Supervised Contrastive Learning for Robust Feature Disentanglement

DGX agent

arXiv:2606.11233v1 Announce Type: new Abstract: Supervised Contrastive Learning (SupCon) has achieved strong performance by explicitly modeling pairwise relationships among samples. However, existing

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

Parameter-Efficient Adapter Tuning for Tabular-Image Multimodal Learning

DGX agent

arXiv:2606.11682v1 Announce Type: new Abstract: Tabular-image multimodal learning aims to improve predictive modeling by jointly using structured tabular attributes and visual data. Although pretraine

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

ParseFixer: An Agentic Framework for Document Parsing via Selective Multimodal Correction

DGX agent

arXiv:2606.11977v1 Announce Type: new Abstract: In this report, we present our third-place solution for the DataMFM Challenge Track 1: Document Parsing. This track requires models to recover structure

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

Phi-Actor-Critic: Steering General-Sum Games to Pareto-Efficient Correlated Equilibria

DGX agent

arXiv:2606.11284v1 Announce Type: cross Abstract: Real-world multi-agent systems, from traffic coordination to resource allocation, are often modeled as general-sum games where individual incentives c

model-releasesarxiv-cs-lg
11 Jun 2026
Model Releases

Robustness of Mixtures of Experts to Feature Noise

DGX agent

arXiv:2601.14792v2 Announce Type: replace Abstract: Despite their practical success, it remains unclear why Mixture of Experts (MoE) models can outperform dense networks beyond sheer parameter scaling

model-releasesarxiv-cs-lg
11 Jun 2026
Safety

SAFER-Nav: Enhancing Safety for Visual Robot Navigation via Segmentation-Aware Fine-Tuning

DGX agent

arXiv:2606.11636v1 Announce Type: new Abstract: Vision-based navigation models, particularly foundation models, generate viable trajectories from RGB observations alone. However, even state-of-the-art

safetyarxiv-cs-ro
11 Jun 2026
Safety

SAGE: Answer-Conditioned Uncertainty Targets for Verbal Uncertainty Alignment

DGX agent

arXiv:2606.11512v1 Announce Type: new Abstract: Large language models increasingly express uncertainty through natural-language statements, yet these expressions often fail to reflect the model's samp

safetyarxiv-cs-cl
11 Jun 2026
Model Releases

Steering the Noise: Turning Random Perturbations into Effective Descent for Memory-Efficient LLM Fine-Tuning

DGX agent

arXiv:2601.04710v2 Announce Type: replace Abstract: Fine-tuning large language models (LLMs) achieves strong performance but is often limited by the memory overhead of backpropagation. Zeroth-order (Z

model-releasesarxiv-cs-cl
11 Jun 2026
Agents

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes

DGX agent

arXiv:2606.11470v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved strong performance across natural language processing tasks, yet reliable reasoning remains an open challenge

agentsarxiv-cs-cl
11 Jun 2026
Model Releases

Using Explainability as a Training-Time Reliability Signal for Efficient ECG Classification

DGX agent

arXiv:2606.12252v1 Announce Type: cross Abstract: Training deep neural networks for clinical time-series analysis is computationally demanding, yet many healthcare settings lack the resources required

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

VL-DINO: Leveraging CLIP Vision-Language Knowledge for Open-Vocabulary Object Detectio

DGX agent

arXiv:2606.11546v1 Announce Type: new Abstract: Vision-language models like CLIP can provide rich semantic priors for open-vocabulary object detection. However, jointly integrating both textual and vi

model-releasesarxiv-cs-cv
11 Jun 2026
Research

Weighted Random Dot Product Graphs

DGX agent

arXiv:2505.03649v4 Announce Type: replace-cross Abstract: Modeling of intricate relational patterns has become a cornerstone of contemporary statistical research and related data science fields. Netwo

researcharxiv-cs-lg
11 Jun 2026
Safety

3SPO: State-Score-Supervised Policy Optimization for LLM Agents

DGX agent

arXiv:2606.09961v1 Announce Type: cross Abstract: Training large language models (LLMs) as autonomous agents via reinforcement learning (RL) has enabled frontier models to achieve superhuman performan

safetyarxiv-cs-ai
10 Jun 2026
Model Releases

5% > 100%: Flatness Preference is All You Need for Multimodal Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2606.10488v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) methods provide a streamlined and efficient tool for adapting large models to domain-specific multimodal downstre

model-releasesarxiv-cs-cv
10 Jun 2026
Safety

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

DGX agent

arXiv:2410.15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical. Direct P

safetyarxiv-cs-ai
10 Jun 2026
Model Releases

A Constrained Natural-Language Interface for Variational Multi-Physics Finite Element Simulations in FEniCS

DGX agent

arXiv:2606.10928v1 Announce Type: cross Abstract: Large language models can reduce the manual effort required to set up finite element simulations, but they introduce reliability risks when generated

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

A History-Aware Visually Grounded Critic for Computer Use Agents

DGX agent

arXiv:2606.11078v1 Announce Type: new Abstract: Various test-time interventions for Computer Use Agents (CUAs), including critic models, have been developed to improve performance through pre-executio

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity

DGX agent

arXiv:2606.11150v1 Announce Type: new Abstract: Large language models (LLMs) are rapidly acquiring capabilities relevant to biological research, from literature synthesis to interpretation of experime

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs

DGX agent

arXiv:2603.14463v2 Announce Type: replace Abstract: Adapting Large Language Models (LLMs) to high-stakes vertical domains like insurance presents a significant challenge: scenarios demand strict adher

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

BadRobot: Jailbreaking Embodied LLM Agents in the Physical World

DGX agent

arXiv:2407.20242v5 Announce Type: replace-cross Abstract: Embodied AI represents systems where AI is integrated into physical entities. Large Language Model (LLM), which exhibits powerful language und

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Beyond Memorization: Distinguishing Between Pattern-Based and Epistemic Reasoning in LLMs Using Epistemic Puzzles

DGX agent

arXiv:2603.21350v2 Announce Type: replace Abstract: Epistemic reasoning requires agents to infer the state of the world from partial observations and information about other agents' knowledge. Prior w

model-releasesarxiv-cs-cl
10 Jun 2026
Research

Blurry Window Attention

DGX agent

arXiv:2606.09862v1 Announce Type: cross Abstract: The Softmax Attention operation in Transformer language models has a quadratic complexity in the sequence length and a growing state size in the form

researcharxiv-cs-ai
10 Jun 2026
Model Releases

CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback

DGX agent

arXiv:2504.02323v4 Announce Type: replace Abstract: Large language models (LLMs) have created new opportunities to assist teachers and support student learning. While researchers have explored various

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

Data-Driven Dynamic Assortment in Online Platforms: Learning about Two Sides

DGX agent

arXiv:2606.11118v1 Announce Type: new Abstract: We study a dynamic assortment problem on a two-sided service platform with incomplete information and heterogeneous customers in a discrete-time setting

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

DB-3DME: From Dataset to Benchmark for Human-aligned Automatic 3D Mesh Evaluation

DGX agent

arXiv:2606.10142v1 Announce Type: new Abstract: Recent advances in 3D generation have led to substantial improvements in realism, controllability, and efficiency, yet the evaluation of 3D assets remai

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

Do VLMs Reason Like Engineers? A Benchmark and a Stage-wise Evaluation

DGX agent

arXiv:2606.10833v1 Announce Type: new Abstract: Vision-Language Models (VLMs) demonstrate strong performance on general multimodal reasoning benchmarks, yet their ability to perform engineering reason

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing

DGX agent

arXiv:2606.10223v1 Announce Type: cross Abstract: Attributing a synthetic utterance to its originating system remains an open challenge: closed-set models fail to reject unseen synthesizers and produc

model-releasesarxiv-cs-ai
10 Jun 2026
← Previous
1…413414415416417…1082
Next →