AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,617
  • Agents7,497
  • Applications5,365
  • Concepts5
  • Hardware1,816
  • Industry6,151
  • Local Ai4,900
  • Model Releases23,593
  • Research19,967
  • Safety13,267
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,617
  • Agents7,497
  • Applications5,365
  • Concepts5
  • Hardware1,816
  • Industry6,151
  • Local Ai4,900
  • Model Releases23,593
  • Research19,967
  • Safety13,267
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

Content type
87,617Total entries
1Added by human
87,616Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
51,512 results
Research

Provable Sparse Inversion and Token Relabel Enhanced One-shot Federated Learning with ViTs

DGX agent

arXiv:2605.10748v1 Announce Type: cross Abstract: One-Shot Federated Learning, where a central server learns a global model in a single communication round, has emerged as a promising paradigm. Howeve

researcharxiv-cs-ai
12 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Quantum Transfer Learning Shows Improved Robustness in Low-Data Regimes

DGX agent

arXiv:2605.09118v1 Announce Type: cross Abstract: Transfer learning under limited data is a challenging setting, where models must adapt to new tasks with minimal supervision. Prior work has primarily

researcharxiv-cs-lg
12 May 2026
Model Releases

Real vs. Semi-Simulated: Rethinking Evaluation for Treatment Effect Estimation

DGX agent

arXiv:2605.10430v1 Announce Type: cross Abstract: Estimating heterogeneous treatment effects with machine learning has attracted substantial attention in both academic research and industrial practice

model-releasesarxiv-cs-ai
12 May 2026
Safety

Reasoning Compression with Mixed-Policy Distillation

DGX agent

arXiv:2605.08776v1 Announce Type: new Abstract: Reasoning-centric large language models (LLMs) achieve strong performance by generating intermediate reasoning trajectories, but often incur excessive t

safetyarxiv-cs-ai
12 May 2026
Model Releases

Robust Server Defense Against Unreliable Clients in One-Shot Fair Collaborative Machine Learning

DGX agent

arXiv:2605.08616v1 Announce Type: new Abstract: Collaborative machine learning (CML) enables multiple clients to train a global model jointly in a data-distributed setting. To address data privacy and

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention

DGX agent

arXiv:2603.22016v2 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) often reach a correct solution before their long Chain-of-Thought trace ends, yet continue with redundant verifi

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Route Before Retrieve: Activating Latent Routing Abilities of LLMs for RAG vs. Long-Context Selection

DGX agent

arXiv:2605.10235v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have expanded the context window to beyond 128K tokens, enabling long-document understanding and multi-s

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

S2FT: Parameter-Efficient Fine-Tuning in Sparse Spectrum Domain

DGX agent

arXiv:2605.08589v1 Announce Type: new Abstract: Parameter Efficient Fine-Tuning (PEFT) is a key technique for adapting a large pretrained model to downstream tasks by fine-tuning only a small number o

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

SayNext-Bench: Why Do LLMs Struggle with Next-Utterance Anticipation?

DGX agent

arXiv:2602.00327v2 Announce Type: replace Abstract: We explore the use of large language models (LLMs) for next-utterance anticipation in human dialogue. Despite recent advances in LLMs demonstrating

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation

DGX agent

arXiv:2605.10187v1 Announce Type: new Abstract: Scientific reasoning is a key aspect of human intelligence, requiring the integration of multimodal inputs, domain expertise, and multi-step inference a

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning

DGX agent

arXiv:2605.09266v1 Announce Type: new Abstract: We introduce SeePhys Pro, a fine-grained modality transfer benchmark that studies whether models preserve the same reasoning capability when critical in

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation

DGX agent

arXiv:2605.10376v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have advanced rapidly in multimodal perception and language understanding, yet it remains unclear whether they can reliabl

model-releasesarxiv-cs-cv
12 May 2026
Research

SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding

DGX agent

arXiv:2605.10453v1 Announce Type: cross Abstract: Speculative decoding speeds up autoregressive generation in Large Language Models (LLMs) through a two-step procedure, where a lightweight draft model

researcharxiv-cs-cl
12 May 2026
Model Releases

Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens

DGX agent

arXiv:2604.02608v2 Announce Type: replace Abstract: Activation steering presupposes that task-relevant behaviors correspond to linear directions in activation space -- directions that should both stee

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution

DGX agent

arXiv:2605.08366v1 Announce Type: new Abstract: We introduce SWE Atlas, a benchmark suite for coding agents spanning three professional software engineering workflows: Codebase Q&A (124 tasks), Test W

model-releasesarxiv-cs-lg
12 May 2026
Research

TAD: Temporal-Aware Trajectory Self-Distillation for Fast and Accurate Diffusion LLM

DGX agent

arXiv:2605.09536v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) offer a promising paradigm for parallel text generation, but in practice they face an accuracy-parallelism tra

researcharxiv-cs-ai
12 May 2026
Agents

Temporal Sampling Frequency Matters: A Capacity-Aware Study of End-to-End Driving Trajectory Prediction

DGX agent

arXiv:2605.10388v1 Announce Type: new Abstract: End to end (E2E) autonomous driving trajectory prediction is often trained with camera frames sampled at the highest available temporal frequency, assum

agentsarxiv-cs-cv
12 May 2026
Model Releases

The Differences Between Direct Alignment Algorithms are a Blur

DGX agent

arXiv:2502.01237v3 Announce Type: replace Abstract: Direct Alignment Algorithms (DAAs) simplify LLM alignment by directly optimizing policies, bypassing reward modeling and RL. While DAAs differ in th

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

To Redact, or not to Redact? A Local LLM Approach to Deliberative Process Privilege Classification

DGX agent

arXiv:2605.10211v1 Announce Type: cross Abstract: Government transparency laws, like the Freedom of Information (FOIA) acts in the United States and United Kingdom, and the Woo (Open Government Act) i

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Tokens-per-Parameter Coverage Is Critical for Robust LLM Scaling Law Extrapolation

DGX agent

arXiv:2605.08541v1 Announce Type: new Abstract: Neural scaling laws approximate a language model's loss as a power-law function of parameter count N and token count D. Following Chinchilla-style compu

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Towards Cross-lingual Values Judgment: A Consensus-Pluralism Perspective

DGX agent

arXiv:2602.17283v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are employed worldwide, existing evaluation paradigms for their multilingual capabilities primarily focus on f

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Two Ways to De-Bias an LLM-as-a-Judge: A Continuous-Score Comparison of Hierarchical Bayesian Calibration and Neural-ODE Score Transport

DGX agent

arXiv:2605.09227v1 Announce Type: new Abstract: [Abridged] Using a Large Language Model (LLM) as an automatic rater (LLM-as-a-judge) is cheap but potentially biased: some judges run lenient, others st

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

UFO: A Unified Flow-Oriented Framework for Robust Continual Graph Learning

DGX agent

arXiv:2605.09862v1 Announce Type: cross Abstract: Graph learning research has increasingly shifted toward continual graph learning (CGL), which better reflects real-world scenarios where graphs evolve

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

UMEDA: Unified Multi-modal Efficient Data Fusion for Privacy-Preserving Graph Federated Learning via Spectral-Gated Attention and Diffusion-Based Operator Alignment

DGX agent

arXiv:2605.08288v1 Announce Type: cross Abstract: Device-free localization trains models from heterogeneous wireless and visual sensors (e.g., Wi-Fi, LiDAR) distributed across edge devices. Federated

model-releasesarxiv-cs-ai
12 May 2026
Safety

What should post-training optimize? A test-time scaling law perspective

DGX agent

arXiv:2605.10716v1 Announce Type: new Abstract: Large language models are increasingly deployed with test-time strategies: sample N responses, score them with a reward model or verifier, and return th

safetyarxiv-cs-lg
12 May 2026
Model Releases

When Attention Beats Fourier: Multi-Scale Transformers for PDE Solving on Irregular Domains

DGX agent

arXiv:2605.08318v1 Announce Type: cross Abstract: We study the problem of architecture selection for deep learning models trained to solve partial differential equations (PDEs), asking when transforme

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

DGX agent

arXiv:2605.10434v1 Announce Type: new Abstract: Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving i

model-releasesarxiv-cs-cv
12 May 2026
Research

A Theory of Online Learning with Autoregressive Chain-of-Thought Reasoning

DGX agent

arXiv:2605.06819v1 Announce Type: new Abstract: Autoregressive generation lies at the heart of the mechanism of large language models. It can be viewed as the repeated application of a next-token gene

researcharxiv-cs-lg
11 May 2026
Safety

Activation Differences Reveal Backdoors: A Comparison of SAE Architectures

DGX agent

arXiv:2605.07324v1 Announce Type: cross Abstract: Backdoor attacks on language models pose a significant threat to AI safety, where models behave normally on most inputs but exhibit harmful behavior w

safetyarxiv-cs-ai
11 May 2026
Model Releases

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

DGX agent

arXiv:2605.06869v1 Announce Type: new Abstract: AI agent research spans a wide spectrum: from RL agents that learn from scratch to foundation model agents that leverage pre-trained knowledge, yet no u

model-releasesarxiv-cs-ai
11 May 2026
Safety

BEAVER: An Efficient Deterministic LLM Verifier

DGX agent

arXiv:2512.05439v2 Announce Type: replace Abstract: As large language models (LLMs) transition from research prototypes to production systems, practitioners often need reliable methods to verify model

safetyarxiv-cs-ai
11 May 2026
Model Releases

Beyond LoRA vs. Full Fine-Tuning: Gradient-Guided Optimizer Routing for LLM Adaptation

DGX agent

arXiv:2605.07111v1 Announce Type: cross Abstract: Recent literature on fine-tuning Large Language Models highlights a fundamental debate. While Full Fine-Tuning (FFT) provides the representational pla

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Beyond the Black Box: Interpretability of Agentic AI Tool Use

DGX agent

arXiv:2605.06890v1 Announce Type: new Abstract: AI agents are promising for high-stakes enterprise workflows, but dependable deployment remains limited because tool-use failures are difficult to diagn

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

CarCrashNet: A Large-Scale Dataset and Hierarchical Neural Solver for Data-Driven Structural Crash Simulation

DGX agent

arXiv:2605.07098v1 Announce Type: new Abstract: Crash simulation is a cornerstone of modern vehicle development because it reduces the need for costly physical prototypes, accelerates safety-driven de

model-releasesarxiv-cs-lg
11 May 2026
Research

CONSIGN: Conformal Segmentation Informed by Spatial Groupings via Decomposition

DGX agent

arXiv:2505.14113v3 Announce Type: replace Abstract: Most machine learning-based image segmentation models produce pixel-wise confidence scores that represent the model's predicted probability for each

researcharxiv-cs-cv
11 May 2026
Model Releases

CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs

DGX agent

arXiv:2605.06115v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs), trained primarily on English-centric data, frequently generate culturally inappropriate or misaligned resp

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

CSR: Infinite-Horizon Real-Time Policies with Massive Cached State Representations

DGX agent

arXiv:2605.07325v1 Announce Type: cross Abstract: Deploying massive large language models (LLMs) as continuous cognitive engines for robotics is bottlenecked by the time-to-first-token (TTFT) latency

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks

DGX agent

arXiv:2603.04676v2 Announce Type: replace-cross Abstract: Multi-image reasoning remains a significant challenge for vision-language models (VLMs). We investigate a previously overlooked phenomenon: du

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Divide and Conquer: Object Co-occurrence Helps Mitigate Simplicity Bias in OOD Detection

DGX agent

arXiv:2605.07821v1 Announce Type: cross Abstract: Out-of-distribution (OOD) detection is crucial for ensuring the reliability of deep learning models. Existing methods mostly focus on regular entangle

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams

DGX agent

arXiv:2605.07299v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) remain primarily reactive, failing to continuously perceive environments or proactively assist users

model-releasesarxiv-cs-ai
11 May 2026
Research

Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning

DGX agent

arXiv:2605.06840v1 Announce Type: new Abstract: Large language models (LLMs), especially reasoning models, generate extended chain-of-thought (CoT) reasoning that often contains explicit deliberation

researcharxiv-cs-ai
11 May 2026
Model Releases

FAME: Forecasting Academic Impact via Continuous-Time Manifold Evolution

DGX agent

arXiv:2605.07208v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used to brainstorm and evaluate research ideas, yet assessing such judgments is fundamentally difficult be

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Gradient Extrapolation-Based Policy Optimization

DGX agent

arXiv:2605.06755v1 Announce Type: cross Abstract: Reinforcement learning is widely used to improve the reasoning ability of large language models, especially when answers can be automatically checked.

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Is Your Prompt Poisoning Code? Defect Induction Rates and Security Mitigation Strategies

DGX agent

arXiv:2510.22944v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have become indispensable for automated code generation, yet the quality and security of their outputs remain a c

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

MedAction: Towards Active Multi-turn Clinical Diagnostic LLMs

DGX agent

arXiv:2605.07305v1 Announce Type: cross Abstract: Most existing LLM diagnoses are evaluated on static, single-turn settings where complete patient information is provided upfront, an oversimplificatio

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

MiniAppBench: Evaluating the Shift from Text to Interactive HTML Responses in LLM-Powered Assistants

DGX agent

arXiv:2603.09652v3 Announce Type: replace Abstract: With the rapid advancement of Large Language Models (LLMs) in code generation, human-AI interaction is evolving from static text responses to dynami

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

MobileDev-Bench: A Benchmark for Issue Resolution in Mobile Application Development

DGX agent

arXiv:2603.24946v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown strong performance on automated software engineering tasks, yet existing benchmarks focus primarily on

model-releasesarxiv-cs-lg
11 May 2026
Safety

Multi-environment Invariance Learning with Missing Data

DGX agent

arXiv:2601.07247v2 Announce Type: replace-cross Abstract: Learning models that can handle distribution shifts is a key challenge in domain generalization. Invariance learning, an approach that focuses

safetyarxiv-cs-lg
11 May 2026
← Previous
1…372373374375376…1074
Next →