AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Safety

RUBEN: Rule-Based Explanations for Retrieval-Augmented LLM Systems

DGX agent

arXiv:2605.10862v1 Announce Type: new Abstract: This paper demonstrates RUBEN, an interactive tool for discovering minimal rules to explain the outputs of retrieval-augmented large language models (LL

safetyarxiv-cs-cl
12 May 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards

DGX agent

arXiv:2605.10899v1 Announce Type: new Abstract: Training deep research agents, namely systems that plan, search, evaluate evidence, and synthesize long-form reports, pushes reinforcement learning beyo

safetyarxiv-cs-cl
12 May 2026
Safety

SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators

DGX agent

arXiv:2605.08334v1 Announce Type: new Abstract: We present SalesSim, a framework and testbed for evaluating the ability of Multimodal Large Language Models (MLLMs) to simulate realistic, persona-drive

safetyarxiv-cs-cl
12 May 2026
Local Ai

Scaling Mobile Agent Systems: From Capability Density to Collective Intelligence

DGX agent

arXiv:2605.08124v1 Announce Type: cross Abstract: Mobile agent systems are emerging as a key paradigm for enabling intelligent applications on edge devices and in AIoT ecosystems. However, their scala

local-aiarxiv-cs-cl
12 May 2026
Research

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models

DGX agent

arXiv:2605.09630v1 Announce Type: new Abstract: Tokenizer-free language models eliminate the tokenizer step of the language modeling pipeline by operating directly on bytes; patch-based variants furth

researcharxiv-cs-cl
12 May 2026
Model Releases

SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness

DGX agent

arXiv:2603.14889v2 Announce Type: replace-cross Abstract: The rapid evolution of end-to-end spoken dialogue systems demands transcending mere textual semantics to incorporate paralinguistic nuances an

model-releasesarxiv-cs-cl
12 May 2026
Agents

SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization

DGX agent

arXiv:2605.08382v1 Announce Type: cross Abstract: LLM coding agents now generate code at an unprecedented scale, yet LLM-generated code introduces cybersecurity vulnerabilities into codebases without

agentsarxiv-cs-cl
12 May 2026
Research

ShifaMind: A Multiplicative Concept Bottleneck for Interpretable ICD-10 Coding

DGX agent

arXiv:2605.08482v1 Announce Type: cross Abstract: Automated ICD-10 coding from clinical discharge summaries requires models that are both accurate on long-tailed multi-label classification tasks and i

researcharxiv-cs-cl
12 May 2026
Agents

SkillMAS: Skill Co-Evolution with LLM-based Multi-Agent System

DGX agent

arXiv:2605.09341v1 Announce Type: cross Abstract: Large language model (LLM) agent systems are increasingly expected to improve after deployment, but existing work often decouples two adaptation targe

agentsarxiv-cs-cl
12 May 2026
Agents

SkillRAE: Agent Skill-Based Context Compilation for Retrieval-Augmented Execution

DGX agent

arXiv:2605.10114v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents (e.g., OpenClaw) increasingly rely on reusable skill libraries to solve artifact-rich tasks such as document-cen

agentsarxiv-cs-cl
12 May 2026
Research

SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding

DGX agent

arXiv:2605.10453v1 Announce Type: cross Abstract: Speculative decoding speeds up autoregressive generation in Large Language Models (LLMs) through a two-step procedure, where a lightweight draft model

researcharxiv-cs-cl
12 May 2026
Model Releases

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

DGX agent

arXiv:2605.09063v1 Announce Type: new Abstract: Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challengi

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection

DGX agent

arXiv:2605.08583v1 Announce Type: new Abstract: Large language models are increasingly used in scientific writing, yet they can fabricate citation-shaped references that appear plausible but fail bibl

model-releasesarxiv-cs-cl
12 May 2026
Research

Sparse Layers are Critical to Scaling Looped Language Models

DGX agent

arXiv:2605.09165v1 Announce Type: cross Abstract: Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundar

researcharxiv-cs-cl
12 May 2026
Tutorials

Sparse Reward Subsystem in Large Language Models

DGX agent

arXiv:2602.00986v2 Announce Type: replace Abstract: Recent studies show that LLM hidden states encode reward-related information, such as answer correctness and model confidence. However, existing app

tutorialsarxiv-cs-cl
12 May 2026
Model Releases

SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?

DGX agent

arXiv:2602.03916v3 Announce Type: replace-cross Abstract: Spatial reasoning is a fundamental aspect of human cognition, yet it remains a major challenge for contemporary vision-language models (VLMs).

model-releasesarxiv-cs-cl
12 May 2026
Research

Spherical Flows for Sampling Categorical Data

DGX agent

arXiv:2605.05629v2 Announce Type: replace-cross Abstract: We study the problem of learning generative models for discrete sequences in a continuous embedding space. Whereas prior approaches typically

researcharxiv-cs-cl
12 May 2026
Research

SSA: Improving Performance With a Better Scoring Function

DGX agent

arXiv:2508.14685v4 Announce Type: replace Abstract: While transformer models exhibit strong in-context learning (ICL) abilities, they often fail to generalize under simple distribution shifts. We anal

researcharxiv-cs-cl
12 May 2026
Model Releases

Statistical Scouting Finds Debate-Safe but Not Debate-Useful Cases: A Matched-Ceiling Study of Open-Weight LLM Reasoning Protocols

DGX agent

arXiv:2605.09618v1 Announce Type: new Abstract: When should a language model answer directly, sample and vote, or engage in multi-agent debate? Recent work shows voting often explains much of the gain

model-releasesarxiv-cs-cl
12 May 2026
Research

Structured Recurrent Mixers for Massively Parallelized Sequence Generation

DGX agent

arXiv:2605.08696v1 Announce Type: new Abstract: Over the last two decades, language modeling has experienced a shift from predominantly recurrent architectures that process tokens sequentially during

researcharxiv-cs-cl
12 May 2026
Model Releases

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data

DGX agent

arXiv:2605.10129v1 Announce Type: new Abstract: Large language models (LLMs) rely on web-scale corpora for pre-training. The noise inherent in these datasets tends to obscure meaningful patterns and u

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems

DGX agent

arXiv:2605.09539v1 Announce Type: new Abstract: Multi-agent systems (MAS) have emerged as a promising paradigm for solving complex tasks. Recent work has explored self-evolving MAS that automatically

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Talk to Your Slides: High-Efficiency Slide Editing via Language-Driven Structured Data Manipulation

DGX agent

arXiv:2505.11604v5 Announce Type: replace Abstract: Editing presentation slides is a frequent yet tedious task, ranging from creative layout design to repetitive text maintenance. While recent GUI-bas

model-releasesarxiv-cs-cl
12 May 2026
Research

Task-Aware Calibration: Provably Optimal Decoding in LLMs

DGX agent

arXiv:2605.10202v1 Announce Type: cross Abstract: LLM decoding often relies on the model's predictive distribution to generate an output. Consequently, misalignment with respect to the true generating

researcharxiv-cs-cl
12 May 2026
Research

TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination

DGX agent

arXiv:2510.22767v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) typically come with a fixed architecture, despite growing evidence that not all layers contribute equally to ever

researcharxiv-cs-cl
12 May 2026
Applications

Temporal Tokenization Strategies for Event Sequence Modeling with Large Language Models

DGX agent

arXiv:2512.13618v3 Announce Type: replace Abstract: Representing continuous time is a critical and under-explored challenge in modeling temporal event sequences with large language models (LLMs). Vari

applicationsarxiv-cs-cl
12 May 2026
Model Releases

Test-Time Speculation

DGX agent

arXiv:2605.09329v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by using a fast draft model to generate tokens and a more accurate target model to verify them. Its perfo

model-releasesarxiv-cs-cl
12 May 2026
Research

The Association of Transformer-based Sentiment Analysis with Symptom Distress and Deterioration in Routine Psychotherapy Care

DGX agent

arXiv:2605.09838v1 Announce Type: new Abstract: Sentiment analysis has been of long-standing interest in psychotherapy research. Recently, the Transformer deep learning architecture has produced text-

researcharxiv-cs-cl
12 May 2026
Research

The Astonishing Ability of Large Language Models to Parse Jabberwockified Language

DGX agent

arXiv:2602.23928v2 Announce Type: replace Abstract: We show that large language models (LLMs) have an astonishing ability to recover meaning from severely degraded English texts. Texts in which conten

researcharxiv-cs-cl
12 May 2026
Model Releases

The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs

DGX agent

arXiv:2605.08737v1 Announce Type: cross Abstract: On-policy distillation (OPD) is widely used for LLM post-training. When pushed with a reward-extrapolation coefficient lambda > 1, the student can lif

model-releasesarxiv-cs-cl
12 May 2026
Research

The Impact of Editorial Intervention on Detecting Native Language Traces

DGX agent

arXiv:2605.10216v1 Announce Type: new Abstract: Native Language Identification (NLI) is the task of determining an author's native language (L1) from their non-native writings. With the advent of huma

researcharxiv-cs-cl
12 May 2026
Model Releases

The Realignment Problem: When Right becomes Wrong in LLMs

DGX agent

arXiv:2511.02623v2 Announce Type: replace Abstract: Post-training alignment of large language models (LLMs) relies on large-scale human annotations guided by policy specifications that change over tim

model-releasesarxiv-cs-cl
12 May 2026
Local Ai

The Truth Lies Somewhere in the Middle (of the Generated Tokens)

DGX agent

arXiv:2605.09969v1 Announce Type: cross Abstract: How should hidden states generated autoregressively be collapsed into a representation that reflects a language model's internal state? Despite tokens

local-aiarxiv-cs-cl
12 May 2026
Agents

Through the Lens of Character: Resolving Modality-Role Interference in Multimodal Role-Playing Agent

DGX agent

arXiv:2605.09443v1 Announce Type: cross Abstract: The advancement of Multimodal Large Language Models (MLLMs) has expanded Role-Playing Agents (RPAs) into visually grounded environments. However, huma

agentsarxiv-cs-cl
12 May 2026
Safety

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning

DGX agent

arXiv:2508.20697v3 Announce Type: replace-cross Abstract: As large language models (LLMs) continue to grow in capability, so do the risks of harmful misuse through fine-tuning. While most prior studie

safetyarxiv-cs-cl
12 May 2026
Agents

ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering

DGX agent

arXiv:2510.20036v2 Announce Type: replace Abstract: Large language model (LLM) agents rely on external tools to solve complex tasks, but real-world toolsets often contain redundant tools with overlapp

agentsarxiv-cs-cl
12 May 2026
Applications

Topological Data Analysis Applications in Natural Language Processing: A Survey

DGX agent

arXiv:2411.10298v5 Announce Type: replace Abstract: The surge of data available on the Internet has driven the adoption of a wide range of computational methods for analyzing and extracting insights f

applicationsarxiv-cs-cl
12 May 2026
Tutorials

Toward Multi-Database Query Reasoning for Text2Cypher

DGX agent

arXiv:2605.10373v1 Announce Type: cross Abstract: Large language models have significantly improved natural language interfaces to databases by translating user questions into executable queries. In p

tutorialsarxiv-cs-cl
12 May 2026
Model Releases

Towards Compact Sign Language Translation: Frame Rate and Model Size Trade-offs

DGX agent

arXiv:2605.09554v1 Announce Type: new Abstract: Sign Language Translation (SLT) converts sign language videos into spoken-language text, bridging communication between Deaf and hearing communities. Cu

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents

DGX agent

arXiv:2605.10832v1 Announce Type: new Abstract: Multimodal deep search requires an agent to solve open-world problems by chaining search, tool use, and visual reasoning over evolving textual and visua

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents

DGX agent

arXiv:2605.09934v1 Announce Type: new Abstract: Multimodal large language models increasingly solve vision-centric tasks by calling external tools for visual inspection, OCR, retrieval, calculation, a

model-releasesarxiv-cs-cl
12 May 2026
Safety

Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning

DGX agent

arXiv:2605.08741v1 Announce Type: new Abstract: Inference-time harnesses substantially improve large language models on complex reasoning tasks. However, the intrinsic capabilities of the underlying m

safetyarxiv-cs-cl
12 May 2026
Model Releases

Two Ways to De-Bias an LLM-as-a-Judge: A Continuous-Score Comparison of Hierarchical Bayesian Calibration and Neural-ODE Score Transport

DGX agent

arXiv:2605.09227v1 Announce Type: new Abstract: [Abridged] Using a Large Language Model (LLM) as an automatic rater (LLM-as-a-judge) is cheap but potentially biased: some judges run lenient, others st

model-releasesarxiv-cs-cl
12 May 2026
Local Ai

UPA: Unsupervised Prompt Agent via Tree-Based Search and Selection

DGX agent

arXiv:2601.23273v2 Announce Type: replace Abstract: Prompt agents have recently emerged as a promising paradigm for automated prompt optimization, framing prompt discovery as a sequential decision-mak

local-aiarxiv-cs-cl
12 May 2026
Model Releases

UserGPT Technical Report

DGX agent

arXiv:2605.08766v1 Announce Type: cross Abstract: Personalized user understanding from large-scale digital traces remains a fundamental challenge. Traditional user profiling methods rely on discrimina

model-releasesarxiv-cs-cl
12 May 2026
Safety

V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning

DGX agent

arXiv:2605.10172v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have achieved remarkable success in general perception, yet complex multi-step visual reasoning remains a per

safetyarxiv-cs-cl
12 May 2026
Research

Verbalized Algorithms: Classical Algorithms are All You Need (Mostly)

DGX agent

arXiv:2509.08150v5 Announce Type: replace Abstract: Reasoning is a fundamentally algorithmic task. Yet current work on LLM-based reasoning relies on free-form generation whose theoretical guarantees (

researcharxiv-cs-cl
12 May 2026
Safety

VISTA: A Generative Egocentric Video Framework for Daily Assistance

DGX agent

arXiv:2605.10579v1 Announce Type: new Abstract: Training AI agents to proactively assist humans in daily activities, from routine household tasks to urgent safety situations, requires large-scale visu

safetyarxiv-cs-cl
12 May 2026
← Previous
1…100101102103104…161
Next →