AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Safety

A Theory of Time-Sensitive Language Generation: Sparse Hallucination Beats Mode Collapse

DGX agent

arXiv:2605.11302v1 Announce Type: cross Abstract: We study language generation in the limit under a global preference ordering on strings, as introduced by Kleinberg and Wei. As in [arXiv:2504.14370,

safetyarxiv-cs-cl
13 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment

DGX agent

arXiv:2605.11398v1 Announce Type: cross Abstract: We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations.

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference

DGX agent

arXiv:2605.11581v1 Announce Type: new Abstract: When large language models (LLMs) serve real-time inference in commercial online advertising systems, end-to-end latency must be strictly bounded to the

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness Predictions

DGX agent

arXiv:2211.03524v2 Announce Type: replace Abstract: Modern Review Helpfulness Prediction systems are dependent upon multiple modalities, typically texts and images. Unfortunately, those contemporary a

model-releasesarxiv-cs-cl
13 May 2026
Safety

Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning

DGX agent

arXiv:2605.11458v1 Announce Type: cross Abstract: On-policy self-distillation has become a strong recipe for LLM reasoning, where a privileged teacher supervises the student's own rollouts while condi

safetyarxiv-cs-cl
13 May 2026
Model Releases

Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty

DGX agent

arXiv:2605.11436v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed on long-horizon tasks in partially observable environments, where they must act while inferring a

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

AgentDisCo: Towards Disentanglement and Collaboration in Open-ended Deep Research Agents

DGX agent

arXiv:2605.11732v1 Announce Type: cross Abstract: In this paper, we present AgentDisCo, a novel Disentangled and Collaborative agentic architecture that formulates deep research as an adversarial opti

model-releasesarxiv-cs-cl
13 May 2026
Agents

AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents

DGX agent

arXiv:2605.11026v1 Announce Type: cross Abstract: Defenses against indirect prompt injection (IPI) in tool-using LLM agents share two structural weaknesses. First, they all attempt to prevent attacks

agentsarxiv-cs-cl
13 May 2026
Model Releases

Allegory of the Cave: Measurement-Grounded Vision-Language Learning

DGX agent

arXiv:2605.11727v1 Announce Type: cross Abstract: Vision-language models typically reason over post-ISP RGB images, although RGB rendering can clip, suppress, or quantize sensor evidence before infere

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

An Empirical Study of Automating Agent Evaluation

DGX agent

arXiv:2605.11378v1 Announce Type: new Abstract: Agent evaluation requires assessing complex multi-step behaviors involving tool use and intermediate reasoning, making it costly and expertise-intensive

model-releasesarxiv-cs-cl
13 May 2026
Safety

Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information

DGX agent

arXiv:2605.11609v1 Announce Type: cross Abstract: On-policy self-distillation, where a student is pulled toward a copy of itself conditioned on privileged context (e.g., a verified solution or feedbac

safetyarxiv-cs-cl
13 May 2026
Safety

Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR

DGX agent

arXiv:2604.04894v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning ability of large language models (LLMs), but it often

safetyarxiv-cs-cl
13 May 2026
Hardware

AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration -- Learning from Cheap, Optimizing Expensive

DGX agent

arXiv:2605.11518v1 Announce Type: cross Abstract: Effectively configuring scalable large language model (LLM) experiments, spanning architecture design, hyperparameter tuning, and beyond, is crucial f

hardwarearxiv-cs-cl
13 May 2026
Model Releases

AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor

DGX agent

arXiv:2601.05752v3 Announce Type: replace Abstract: We introduce AutoMonitor-Bench, the first benchmark designed to systematically evaluate the reliability of LLM-based misbehavior monitors across div

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

BEExformer: A Fast Inferencing Binarized Transformer with Early Exits

DGX agent

arXiv:2412.05225v3 Announce Type: replace Abstract: Large Language Models (LLMs) based on transformers achieve cutting-edge results on a variety of applications. However, their enormous size and proce

model-releasesarxiv-cs-cl
13 May 2026
Research

BitLM: Unlocking Multi-Token Language Generation with Bitwise Continuous Diffusion

DGX agent

arXiv:2605.11577v1 Announce Type: new Abstract: Autoregressive language models generate text one token at a time, yet natural language is inherently structured in multi-token units, including phrases,

researcharxiv-cs-cl
13 May 2026
Model Releases

Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents

DGX agent

arXiv:2508.07642v3 Announce Type: replace-cross Abstract: Vision-and-Language Navigation (VLN) poses significant challenges for agents to interpret natural language instructions and navigate complex 3

model-releasesarxiv-cs-cl
13 May 2026
Safety

Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection

DGX agent

arXiv:2605.11442v1 Announce Type: cross Abstract: Large Language Model (LLM) agents have emerged as key intermediaries, orchestrating complex interactions between human users and a wide range of digit

safetyarxiv-cs-cl
13 May 2026
Model Releases

Caraman at SemEval-2026 Task 8: Three-Stage Multi-Turn Retrieval with Query Rewriting, Hybrid Search, and Cross-Encoder Reranking

DGX agent

arXiv:2605.12028v1 Announce Type: new Abstract: We describe our system for SemEval-2026 Task 8 (MTRAGEval), participating in Task A (Retrieval) across four English-language domains. Our approach emplo

model-releasesarxiv-cs-cl
13 May 2026
Safety

Characterizing the Robustness of Black-Box LLM Planners Under Perturbed Observations with Adaptive Stress Testing

DGX agent

arXiv:2505.05665v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently demonstrated success in decision-making tasks including planning, control, and prediction, but thei

safetyarxiv-cs-cl
13 May 2026
Model Releases

Checkup2Action: A Multimodal Clinical Check-up Report Dataset for Patient-Oriented Action Card Generation

DGX agent

arXiv:2605.11533v1 Announce Type: new Abstract: Clinical check-up reports are multimodal documents that combine page layouts, tables, numerical biomarkers, abnormality flags, imaging findings, and dom

model-releasesarxiv-cs-cl
13 May 2026
Research

Choosing features for classifying multiword expressions

DGX agent

arXiv:2605.11779v1 Announce Type: new Abstract: Multiword expressions (MWEs) are a heterogeneous set with a glaring need for classifications. Designing a satisfactory classification involves choosing

researcharxiv-cs-cl
13 May 2026
Model Releases

ClinicalBench: Stress-Testing Assertion-Aware Retrieval for Cross-Admission Clinical QA on MIMIC-IV

DGX agent

arXiv:2605.11143v1 Announce Type: new Abstract: Reasoning benchmarks measure clinical performance on clean inputs. We evaluate the step before reasoning: retrieval over real EHR notes, where negation,

model-releasesarxiv-cs-cl
13 May 2026
Applications

Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner

DGX agent

arXiv:2510.03206v2 Announce Type: replace-cross Abstract: Diffusion language models, especially masked discrete diffusion models, have achieved great success recently. While there are some theoretical

applicationsarxiv-cs-cl
13 May 2026
Safety

Combining On-Policy Optimization and Distillation for Long-Context Reasoning in Large Language Models

DGX agent

arXiv:2605.12227v1 Announce Type: new Abstract: Adapting large language models (LLMs) to long-context tasks requires post-training methods that remain accurate and coherent over thousands of tokens. E

safetyarxiv-cs-cl
13 May 2026
Applications

Concordance Comparison as a Means of Assembling Local Grammars

DGX agent

arXiv:2605.11862v1 Announce Type: new Abstract: Named Entity Recognition for person names is an important but non-trivial task in information extraction. This article uses a tool that compares the con

applicationsarxiv-cs-cl
13 May 2026
Research

Context Convergence Improves Answering Inferential Questions

DGX agent

arXiv:2605.12370v1 Announce Type: new Abstract: While Large Language Models (LLMs) are widely used in open-domain Question Answering (QA), their ability to handle inferential questions-where answers m

researcharxiv-cs-cl
13 May 2026
Safety

Controllable User Simulation

DGX agent

arXiv:2605.11519v1 Announce Type: cross Abstract: Using offline datasets to evaluate conversational agents often fails to cover rare scenarios or to support testing new policies. This has motivated th

safetyarxiv-cs-cl
13 May 2026
Model Releases

Correcting Selection Bias in Sparse User Feedback for Large Language Model Quality Estimation: A Multi-Agent Hierarchical Bayesian Approach

DGX agent

arXiv:2605.12177v1 Announce Type: new Abstract: [Abridged] Production LLM deployments receive feedback from a non-random fraction of users: thumbs sit mostly in the tails of the satisfaction distribut

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification

DGX agent

arXiv:2603.28488v2 Announce Type: replace Abstract: Large language models (LLMs) remain unreliable for high-stakes claim verification due to hallucinations and shallow reasoning. While retrieval-augme

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Decomposing Evolutionary Mixture-of-LoRA Architectures: The Routing Lever, the Lifecycle Penalty, and a Substrate-Conditional Boundary

DGX agent

arXiv:2605.11153v1 Announce Type: new Abstract: We decompose an evolutionary mixture-of-LoRA system on a from-scratch ~150M-parameter widened-D substrate (D=1536, V=32000; D/V approx 0.048; the 'widen

model-releasesarxiv-cs-cl
13 May 2026
Agents

Deep Reasoning in General Purpose Agents via Structured Meta-Cognition

DGX agent

arXiv:2605.11388v1 Announce Type: new Abstract: Humans intuitively solve complex problems by flexibly shifting among reasoning modes: they plan, execute, revise intermediate goals, resolve ambiguity t

agentsarxiv-cs-cl
13 May 2026
Local Ai

DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding

DGX agent

arXiv:2312.02549v2 Announce Type: replace-cross Abstract: Temporal Language Grounding seeks to localize video moments that semantically correspond to a natural language query. Recent advances employ t

local-aiarxiv-cs-cl
13 May 2026
Research

Demystifying When Pruning Works via Representation Hierarchies

DGX agent

arXiv:2603.24652v3 Announce Type: replace Abstract: Network pruning, which removes less important parameters or architectures, is often expected to improve efficiency while preserving performance. How

researcharxiv-cs-cl
13 May 2026
Research

Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models

DGX agent

arXiv:2605.12138v1 Announce Type: cross Abstract: Generating realistic and user-preferred advertisements is a key challenge in e-commerce. Existing approaches utilize multiple independent models drive

researcharxiv-cs-cl
13 May 2026
Model Releases

Detecting Data Contamination in LLMs via In-Context Learning

DGX agent

arXiv:2510.27055v2 Announce Type: replace Abstract: We present Contamination Detection via Context (CoDeC), a practical and accurate method to detect and quantify training data contamination in large

model-releasesarxiv-cs-cl
13 May 2026
Research

Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG)

DGX agent

arXiv:2510.06719v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by grounding them in external knowledge. However, its application i

researcharxiv-cs-cl
13 May 2026
Model Releases

DiffScore: Text Evaluation Beyond Autoregressive Likelihood

DGX agent

arXiv:2605.11601v1 Announce Type: new Abstract: Autoregressive language models are widely used for text evaluation, however, their left-to-right factorization introduces positional bias, i.e., early t

model-releasesarxiv-cs-cl
13 May 2026
Safety

Diffusion-State Policy Optimization for Masked Diffusion Language Models

DGX agent

arXiv:2602.06462v3 Announce Type: replace Abstract: Masked diffusion language models generate text through iterative masked-token filling, but terminal-only rewards on final completions provide coarse

safetyarxiv-cs-cl
13 May 2026
Model Releases

Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics

DGX agent

arXiv:2605.12178v1 Announce Type: cross Abstract: World models enable agents to anticipate the effects of their actions by internalizing environment dynamics. In enterprise systems, however, these dyn

model-releasesarxiv-cs-cl
13 May 2026
Research

Do Language Models Encode Knowledge of Linguistic Constraint Violations?

DGX agent

arXiv:2605.12055v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong linguistic performance, yet their internal mechanisms for producing these predictions remain unclear. We inv

researcharxiv-cs-cl
13 May 2026
Research

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation

DGX agent

arXiv:2510.04265v4 Announce Type: replace-cross Abstract: Pass@k is widely used to report the reasoning performance of LLMs, but it often produces unstable and potentially misleading rankings, especia

researcharxiv-cs-cl
13 May 2026
Agents

DreamAvoid: Critical-Phase Test-Time Dreaming to Avoid Failures in VLA Policies

DGX agent

arXiv:2605.11750v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are often brittle in fine-grained manipulation, where minor action errors during the critical phases can rapidly e

agentsarxiv-cs-cl
13 May 2026
Applications

Efficient LLM-based Advertising via Model Compression and Parallel Verification

DGX agent

arXiv:2605.11582v1 Announce Type: new Abstract: Large language models (LLMs) have shown remarkable potential in advertising scenarios such as ad creative generation and targeted advertising. However,

applicationsarxiv-cs-cl
13 May 2026
Safety

Enhancing Multilingual Counterfactual Generation through Alignment-as-Preference Optimization

DGX agent

arXiv:2605.11632v1 Announce Type: new Abstract: Self-generated counterfactual explanations (SCEs) are minimally modified inputs (minimality) generated by large language models (LLMs) that flip their o

safetyarxiv-cs-cl
13 May 2026
Safety

Enhancing Target-Guided Proactive Dialogue Systems via Conversational Scenario Modeling and Intent-Keyword Bridging

DGX agent

arXiv:2605.11964v1 Announce Type: new Abstract: A target-guided proactive dialogue system aims to steer conversations proactively toward pre-defined targets, such as designated keywords or specific to

safetyarxiv-cs-cl
13 May 2026
Research

Enriching and Controlling Global Semantics for Text Summarization

DGX agent

arXiv:2109.10616v2 Announce Type: replace Abstract: Recently, Transformer-based models have been proven effective in the abstractive summarization task by creating fluent and informative summaries. Ne

researcharxiv-cs-cl
13 May 2026
Safety

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control

DGX agent

arXiv:2605.11775v1 Announce Type: cross Abstract: Policy entropy has emerged as a fundamental measure for understanding and controlling exploration in reinforcement learning with verifiable rewards (R

safetyarxiv-cs-cl
13 May 2026
← Previous
1…9394959697…161
Next →