AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
15 May 2026

Uncertainty Quantification for Large Language Diffusion Models

ResearchDGX agent

arXiv:2605.14570v1 Announce Type: new Abstract: Large Language Diffusion Models (LLDMs) are emerging as an alternative to autoregressive models, offering faster inference through higher parallelism. S

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use

Model ReleasesDGX agent

arXiv:2605.13989v1 Announce Type: new Abstract: We present VectraYX-Nano, a 41.95M-parameter decoder-only language model trained from scratch in Spanish for cybersecurity, with a Latin-American focus

What Do AI Agents Talk About? Discourse and Architectural Constraints in the First AI-Only Social Network

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2603.07880v5 Announce Type: replace Abstract: Moltbook is the first large-scale social network built for autonomous AI agent-to-agent interaction. Early studies on Moltbook have interpreted its

What Makes Words Hard? Sakura at BEA 2026 Shared Task on Vocabulary Difficulty Prediction

ApplicationsDGX agent

arXiv:2605.14257v1 Announce Type: new Abstract: We describe two types of models for vocabulary difficulty prediction: a high-accuracy black-box model, which achieved the top shared task result in the

When Evidence Conflicts: Uncertainty and Order Effects in Retrieval-Augmented Biomedical Question Answering

ResearchDGX agent

arXiv:2605.14115v1 Announce Type: new Abstract: Biomedical retrieval-augmented large language models (LLMs) often face evidence that is incomplete, misleading, or internally contradictory, yet evaluat

13 May 2026

A categorical error sensitivity index (ISEC): A preventive ordinal decision-support measure for irrecoverable errors in manual data entry systems

ResearchDGX agent

arXiv:2605.12328v1 Announce Type: new Abstract: Data entry systems remain structurally vulnerable to categorical misclassifications, particularly in small and medium sized enterprises (SMEs). When nom

A Causal Language Modeling Detour Improves Encoder Continued Pretraining

ResearchDGX agent

arXiv:2605.12438v1 Announce Type: new Abstract: When adapting an encoder to a new domain, the standard approach is to continue training with Masked Language Modeling (MLM). We show that temporarily sw

A Comparative Study of Controlled Text Generation Systems Using Level-Playing-Field Evaluation Principles

ResearchDGX agent

arXiv:2605.12395v1 Announce Type: new Abstract: Background: Many different approaches to controlled text generation (CTG) have been proposed over recent years, but it is difficult to get a clear pictu

A Formal Comparison Between Chain of Thought and Latent Thought

ResearchDGX agent

arXiv:2509.25239v3 Announce Type: replace-cross Abstract: Chain of thought (CoT) elicits reasoning in large language models by explicitly generating intermediate tokens. In contrast, latent thought re

A Study on Hidden Layer Distillation for Large Language Model Pre-Training

Model ReleasesDGX agent

arXiv:2605.11513v1 Announce Type: new Abstract: Knowledge Distillation (KD) is a critical tool for training Large Language Models (LLMs), yet the majority of research focuses on approaches that rely s

A Survey of On-Policy Distillation for Large Language Models

SafetyDGX agent

arXiv:2604.00626v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) continue to grow in both capability and cost, transferring frontier capabilities into smaller, deployable stud

A Theoretical Analysis of Why Masked Diffusion Models Mitigate the Reversal Curse

Model ReleasesDGX agent

arXiv:2602.02133v2 Announce Type: replace-cross Abstract: Autoregressive language models (ARMs) suffer from the reversal curse: after learning ''A is B,'' they often fail on the reverse query ''B is A

A Theory of Time-Sensitive Language Generation: Sparse Hallucination Beats Mode Collapse

SafetyDGX agent

arXiv:2605.11302v1 Announce Type: cross Abstract: We study language generation in the limit under a global preference ordering on strings, as introduced by Kleinberg and Wei. As in [arXiv:2504.14370,

AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment

Model ReleasesDGX agent

arXiv:2605.11398v1 Announce Type: cross Abstract: We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations.

Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference

Model ReleasesDGX agent

arXiv:2605.11581v1 Announce Type: new Abstract: When large language models (LLMs) serve real-time inference in commercial online advertising systems, end-to-end latency must be strictly bounded to the

Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness Predictions

Model ReleasesDGX agent

arXiv:2211.03524v2 Announce Type: replace Abstract: Modern Review Helpfulness Prediction systems are dependent upon multiple modalities, typically texts and images. Unfortunately, those contemporary a

Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning

SafetyDGX agent

arXiv:2605.11458v1 Announce Type: cross Abstract: On-policy self-distillation has become a strong recipe for LLM reasoning, where a privileged teacher supervises the student's own rollouts while condi

Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty

Model ReleasesDGX agent

arXiv:2605.11436v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed on long-horizon tasks in partially observable environments, where they must act while inferring a

AgentDisCo: Towards Disentanglement and Collaboration in Open-ended Deep Research Agents

Model ReleasesDGX agent

arXiv:2605.11732v1 Announce Type: cross Abstract: In this paper, we present AgentDisCo, a novel Disentangled and Collaborative agentic architecture that formulates deep research as an adversarial opti

AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents

AgentsDGX agent

arXiv:2605.11026v1 Announce Type: cross Abstract: Defenses against indirect prompt injection (IPI) in tool-using LLM agents share two structural weaknesses. First, they all attempt to prevent attacks

Allegory of the Cave: Measurement-Grounded Vision-Language Learning

Model ReleasesDGX agent

arXiv:2605.11727v1 Announce Type: cross Abstract: Vision-language models typically reason over post-ISP RGB images, although RGB rendering can clip, suppress, or quantize sensor evidence before infere

An Empirical Study of Automating Agent Evaluation

Model ReleasesDGX agent

arXiv:2605.11378v1 Announce Type: new Abstract: Agent evaluation requires assessing complex multi-step behaviors involving tool use and intermediate reasoning, making it costly and expertise-intensive

Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information

SafetyDGX agent

arXiv:2605.11609v1 Announce Type: cross Abstract: On-policy self-distillation, where a student is pulled toward a copy of itself conditioned on privileged context (e.g., a verified solution or feedbac

Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR

SafetyDGX agent

arXiv:2604.04894v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning ability of large language models (LLMs), but it often

AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration -- Learning from Cheap, Optimizing Expensive

HardwareDGX agent

arXiv:2605.11518v1 Announce Type: cross Abstract: Effectively configuring scalable large language model (LLM) experiments, spanning architecture design, hyperparameter tuning, and beyond, is crucial f

AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor

Model ReleasesDGX agent

arXiv:2601.05752v3 Announce Type: replace Abstract: We introduce AutoMonitor-Bench, the first benchmark designed to systematically evaluate the reliability of LLM-based misbehavior monitors across div

BEExformer: A Fast Inferencing Binarized Transformer with Early Exits

Model ReleasesDGX agent

arXiv:2412.05225v3 Announce Type: replace Abstract: Large Language Models (LLMs) based on transformers achieve cutting-edge results on a variety of applications. However, their enormous size and proce

BitLM: Unlocking Multi-Token Language Generation with Bitwise Continuous Diffusion

ResearchDGX agent

arXiv:2605.11577v1 Announce Type: new Abstract: Autoregressive language models generate text one token at a time, yet natural language is inherently structured in multi-token units, including phrases,

Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents

Model ReleasesDGX agent

arXiv:2508.07642v3 Announce Type: replace-cross Abstract: Vision-and-Language Navigation (VLN) poses significant challenges for agents to interpret natural language instructions and navigate complex 3

Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection

SafetyDGX agent

arXiv:2605.11442v1 Announce Type: cross Abstract: Large Language Model (LLM) agents have emerged as key intermediaries, orchestrating complex interactions between human users and a wide range of digit

Caraman at SemEval-2026 Task 8: Three-Stage Multi-Turn Retrieval with Query Rewriting, Hybrid Search, and Cross-Encoder Reranking

Model ReleasesDGX agent

arXiv:2605.12028v1 Announce Type: new Abstract: We describe our system for SemEval-2026 Task 8 (MTRAGEval), participating in Task A (Retrieval) across four English-language domains. Our approach emplo

Characterizing the Robustness of Black-Box LLM Planners Under Perturbed Observations with Adaptive Stress Testing

SafetyDGX agent

arXiv:2505.05665v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently demonstrated success in decision-making tasks including planning, control, and prediction, but thei

Checkup2Action: A Multimodal Clinical Check-up Report Dataset for Patient-Oriented Action Card Generation

Model ReleasesDGX agent

arXiv:2605.11533v1 Announce Type: new Abstract: Clinical check-up reports are multimodal documents that combine page layouts, tables, numerical biomarkers, abnormality flags, imaging findings, and dom

Choosing features for classifying multiword expressions

ResearchDGX agent

arXiv:2605.11779v1 Announce Type: new Abstract: Multiword expressions (MWEs) are a heterogeneous set with a glaring need for classifications. Designing a satisfactory classification involves choosing

ClinicalBench: Stress-Testing Assertion-Aware Retrieval for Cross-Admission Clinical QA on MIMIC-IV

Model ReleasesDGX agent

arXiv:2605.11143v1 Announce Type: new Abstract: Reasoning benchmarks measure clinical performance on clean inputs. We evaluate the step before reasoning: retrieval over real EHR notes, where negation,

Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner

ApplicationsDGX agent

arXiv:2510.03206v2 Announce Type: replace-cross Abstract: Diffusion language models, especially masked discrete diffusion models, have achieved great success recently. While there are some theoretical

Combining On-Policy Optimization and Distillation for Long-Context Reasoning in Large Language Models

SafetyDGX agent

arXiv:2605.12227v1 Announce Type: new Abstract: Adapting large language models (LLMs) to long-context tasks requires post-training methods that remain accurate and coherent over thousands of tokens. E

Concordance Comparison as a Means of Assembling Local Grammars

ApplicationsDGX agent

arXiv:2605.11862v1 Announce Type: new Abstract: Named Entity Recognition for person names is an important but non-trivial task in information extraction. This article uses a tool that compares the con

Context Convergence Improves Answering Inferential Questions

ResearchDGX agent

arXiv:2605.12370v1 Announce Type: new Abstract: While Large Language Models (LLMs) are widely used in open-domain Question Answering (QA), their ability to handle inferential questions-where answers m

Controllable User Simulation

SafetyDGX agent

arXiv:2605.11519v1 Announce Type: cross Abstract: Using offline datasets to evaluate conversational agents often fails to cover rare scenarios or to support testing new policies. This has motivated th

Correcting Selection Bias in Sparse User Feedback for Large Language Model Quality Estimation: A Multi-Agent Hierarchical Bayesian Approach

Model ReleasesDGX agent

arXiv:2605.12177v1 Announce Type: new Abstract: [Abridged] Production LLM deployments receive feedback from a non-random fraction of users: thumbs sit mostly in the tails of the satisfaction distribut

Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification

Model ReleasesDGX agent

arXiv:2603.28488v2 Announce Type: replace Abstract: Large language models (LLMs) remain unreliable for high-stakes claim verification due to hallucinations and shallow reasoning. While retrieval-augme

Decomposing Evolutionary Mixture-of-LoRA Architectures: The Routing Lever, the Lifecycle Penalty, and a Substrate-Conditional Boundary

Model ReleasesDGX agent

arXiv:2605.11153v1 Announce Type: new Abstract: We decompose an evolutionary mixture-of-LoRA system on a from-scratch ~150M-parameter widened-D substrate (D=1536, V=32000; D/V approx 0.048; the 'widen

Deep Reasoning in General Purpose Agents via Structured Meta-Cognition

AgentsDGX agent

arXiv:2605.11388v1 Announce Type: new Abstract: Humans intuitively solve complex problems by flexibly shifting among reasoning modes: they plan, execute, revise intermediate goals, resolve ambiguity t

DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding

Local AiDGX agent

arXiv:2312.02549v2 Announce Type: replace-cross Abstract: Temporal Language Grounding seeks to localize video moments that semantically correspond to a natural language query. Recent advances employ t

Demystifying When Pruning Works via Representation Hierarchies

ResearchDGX agent

arXiv:2603.24652v3 Announce Type: replace Abstract: Network pruning, which removes less important parameters or architectures, is often expected to improve efficiency while preserving performance. How

Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models

ResearchDGX agent

arXiv:2605.12138v1 Announce Type: cross Abstract: Generating realistic and user-preferred advertisements is a key challenge in e-commerce. Existing approaches utilize multiple independent models drive

Detecting Data Contamination in LLMs via In-Context Learning

Model ReleasesDGX agent

arXiv:2510.27055v2 Announce Type: replace Abstract: We present Contamination Detection via Context (CoDeC), a practical and accurate method to detect and quantify training data contamination in large

Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG)

ResearchDGX agent

arXiv:2510.06719v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by grounding them in external knowledge. However, its application i

DiffScore: Text Evaluation Beyond Autoregressive Likelihood

Model ReleasesDGX agent

arXiv:2605.11601v1 Announce Type: new Abstract: Autoregressive language models are widely used for text evaluation, however, their left-to-right factorization introduces positional bias, i.e., early t

Diffusion-State Policy Optimization for Masked Diffusion Language Models

SafetyDGX agent

arXiv:2602.06462v3 Announce Type: replace Abstract: Masked diffusion language models generate text through iterative masked-token filling, but terminal-only rewards on final completions provide coarse

Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics

Model ReleasesDGX agent

arXiv:2605.12178v1 Announce Type: cross Abstract: World models enable agents to anticipate the effects of their actions by internalizing environment dynamics. In enterprise systems, however, these dyn

Do Language Models Encode Knowledge of Linguistic Constraint Violations?

ResearchDGX agent

arXiv:2605.12055v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong linguistic performance, yet their internal mechanisms for producing these predictions remain unclear. We inv

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation

ResearchDGX agent

arXiv:2510.04265v4 Announce Type: replace-cross Abstract: Pass@k is widely used to report the reasoning performance of LLMs, but it often produces unstable and potentially misleading rankings, especia

DreamAvoid: Critical-Phase Test-Time Dreaming to Avoid Failures in VLA Policies

AgentsDGX agent

arXiv:2605.11750v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are often brittle in fine-grained manipulation, where minor action errors during the critical phases can rapidly e

Efficient LLM-based Advertising via Model Compression and Parallel Verification

ApplicationsDGX agent

arXiv:2605.11582v1 Announce Type: new Abstract: Large language models (LLMs) have shown remarkable potential in advertising scenarios such as ad creative generation and targeted advertising. However,

Enhancing Multilingual Counterfactual Generation through Alignment-as-Preference Optimization

SafetyDGX agent

arXiv:2605.11632v1 Announce Type: new Abstract: Self-generated counterfactual explanations (SCEs) are minimally modified inputs (minimality) generated by large language models (LLMs) that flip their o

Enhancing Target-Guided Proactive Dialogue Systems via Conversational Scenario Modeling and Intent-Keyword Bridging

SafetyDGX agent

arXiv:2605.11964v1 Announce Type: new Abstract: A target-guided proactive dialogue system aims to steer conversations proactively toward pre-defined targets, such as designated keywords or specific to

Enriching and Controlling Global Semantics for Text Summarization

ResearchDGX agent

arXiv:2109.10616v2 Announce Type: replace Abstract: Recently, Transformer-based models have been proven effective in the abstractive summarization task by creating fluent and informative summaries. Ne

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control

SafetyDGX agent

arXiv:2605.11775v1 Announce Type: cross Abstract: Policy entropy has emerged as a fundamental measure for understanding and controlling exploration in reinforcement learning with verifiable rewards (R

← Previous
1…7475767778…129
Next →