AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
15 Apr 2026

Agentic Insight Generation in VSM Simulations

AgentsDGX agent

arXiv:2604.12421v1 Announce Type: new Abstract: Extracting actionable insights from complex value stream map simulations can be challenging, time-consuming, and error-prone. Recent advances in large l

AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs

AgentsDGX agent

arXiv:2604.12179v1 Announce Type: new Abstract: Recent advancements in Large Language Models (LLMs) have improved their ability to process extended conversational contexts, yet fine-tuning and evaluat

AGSC: Adaptive Granularity and Semantic Clustering for Uncertainty Quantification in Long-text Generation

ResearchDGX agent

arXiv:2604.06812v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated impressive capabilities in long-form generation, yet their application is hindered by the hallucinati


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

AlphaEval: Evaluating Agents in Production

Model ReleasesDGX agent

arXiv:2604.12162v1 Announce Type: new Abstract: The rapid deployment of AI agents in commercial settings has outpaced the development of evaluation methodologies that reflect production realities. Exi

[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic

ResearchDGX agent

arXiv:2602.18899v3 Announce Type: replace-cross Abstract: Self-supervised speech models (S3Ms) are known to encode rich phonetic information, yet how this information is structured remains underexplor

Beyond Majority Voting: Efficient Best-Of-N with Radial Consensus Score

AgentsDGX agent

arXiv:2604.12196v1 Announce Type: new Abstract: Large language models (LLMs) frequently generate multiple candidate responses for a given prompt, yet selecting the most reliable one remains challengin

Beyond Single-Dimension Novelty: How Combinations of Theory, Method, and Results-based Novelty Shape Scientific Impact

Model ReleasesDGX agent

arXiv:2604.12471v1 Announce Type: cross Abstract: Scientific novelty drives advances at the research frontier, yet it is also associated with heightened uncertainty and potential resistance from incum

Beyond Transcription: Unified Audio Schema for Perception-Aware AudioLLMs

SafetyDGX agent

arXiv:2604.12506v1 Announce Type: new Abstract: Recent Audio Large Language Models (AudioLLMs) exhibit a striking performance inversion: while excelling at complex reasoning tasks, they consistently u

Calibrated Confidence Estimation for Tabular Question Answering

ResearchDGX agent

arXiv:2604.12491v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed for tabular question answering, yet calibration on structured data is largely unstudied. This pap

Characterizing Human Semantic Navigation in Concept Production as Trajectories in Embedding Space

ApplicationsDGX agent

arXiv:2602.05971v2 Announce Type: replace Abstract: Semantic representations can be framed as a structured, dynamic knowledge space through which humans navigate to retrieve and manipulate meaning. To

CLEAR: Cross-Lingual Enhancement in Alignment via Reverse-training

SafetyDGX agent

arXiv:2604.05821v2 Announce Type: replace Abstract: Existing multilingual embedding models often encounter challenges in cross-lingual scenarios due to imbalanced linguistic resources and less conside

CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation

Model ReleasesDGX agent

arXiv:2604.12268v1 Announce Type: cross Abstract: Large language models (LLMs) can generate code from natural language, but the extent to which they capture intended program behavior remains unclear.

CoG: Controllable Graph Reasoning via Relational Blueprints and Failure-Aware Refinement over Knowledge Graphs

ResearchDGX agent

arXiv:2601.11047v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable reasoning capabilities but often grapple with reliability challenges like hallucinations.

Compiling Activation Steering into Weights via Null-Space Constraints for Stealthy Backdoors

SafetyDGX agent

arXiv:2604.12359v1 Announce Type: cross Abstract: Safety-aligned large language models (LLMs) are increasingly deployed in real-world pipelines, yet this deployment also enlarges the supply-chain atta

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems

Model ReleasesDGX agent

arXiv:2604.12312v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as task-oriented agents in enterprise environments, ensuring their strict adherence to complex

ContextLens: Modeling Imperfect Privacy and Safety Context for Legal Compliance

SafetyDGX agent

arXiv:2604.12308v1 Announce Type: new Abstract: Individuals' concerns about data privacy and AI safety are highly contextualized and extend beyond sensitive patterns. Addressing these issues requires

CoRoVA: Compressed Representations for Vector-Augmented Code Completion

ResearchDGX agent

arXiv:2510.19644v2 Announce Type: replace Abstract: Retrieval-augmented generation has emerged as one of the most effective approaches for code completion enhancement, especially when repository-level

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task

ResearchDGX agent

arXiv:2604.12426v1 Announce Type: cross Abstract: We investigate whether transformers use their depth adaptively across tasks of increasing difficulty. Using a controlled multi-hop relational reasonin

Do VLMs Truly 'Read' Candlesticks? A Multi-Scale Benchmark for Visual Stock Price Forecasting

Model ReleasesDGX agent

arXiv:2604.12659v1 Announce Type: cross Abstract: Vision-language models(VLMs) are increasingly applied to visual stock price forecasting, yet existing benchmarks inadequately evaluate their understan

E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning

ResearchDGX agent

arXiv:2409.06679v3 Announce Type: replace Abstract: Processing long contexts is increasingly important for Large Language Models (LLMs) in tasks like multi-turn dialogues, code generation, and documen

Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects

ResearchDGX agent

arXiv:2604.05546v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) enable sophisticated reasoning over images and videos, yet their inference is hindered by a systemic efficiency

Empirical Evaluation of PDF Parsing and Chunking for Financial Question Answering with RAG

Model ReleasesDGX agent

arXiv:2604.12047v1 Announce Type: new Abstract: PDF files are primarily intended for human reading rather than automated processing. In addition, the heterogeneous content of PDFs, such as text, table

Enhance-then-Balance Modality Collaboration for Robust Multimodal Sentiment Analysis

ResearchDGX agent

arXiv:2604.12518v1 Announce Type: new Abstract: Multimodal sentiment analysis (MSA) integrates heterogeneous text, audio, and visual signals to infer human emotions. While recent approaches leverage c

Enhancing Agentic Textual Graph Retrieval with Synthetic Stepwise Supervision

SafetyDGX agent

arXiv:2510.03323v2 Announce Type: replace Abstract: Integrating textual graphs into Large Language Models (LLMs) is promising for complex graph-based QA. However, a key bottleneck is retrieving inform

Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors

ApplicationsDGX agent

arXiv:2510.09536v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in multilingual, real-world applications with user inputs -- naturally introducing typographi

EvoSpark: Endogenous Interactive Agent Societies for Unified Long-Horizon Narrative Evolution

SafetyDGX agent

arXiv:2604.12776v1 Announce Type: new Abstract: Realizing endogenous narrative evolution in LLM-based multi-agent systems is hindered by the inherent stochasticity of generative emergence. In particul

FABLE: Fine-grained Fact Anchoring for Unstructured Model Editing

Model ReleasesDGX agent

arXiv:2604.12559v1 Announce Type: new Abstract: Unstructured model editing aims to update models with real-world text, yet existing methods often memorize text holistically without reliable fine-grain

From Imitation to Discrimination: Progressive Curriculum Learning for Robust Web Navigation

Model ReleasesDGX agent

arXiv:2604.12666v1 Announce Type: cross Abstract: Text-based web agents offer computational efficiency for autonomous web navigation, yet developing robust agents remains challenging due to the noisy

From Myopic Selection to Long-Horizon Awareness: Sequential LLM Routing for Multi-Turn Dialogue

SafetyDGX agent

arXiv:2604.12385v1 Announce Type: new Abstract: Multi-turn dialogue is the predominant form of interaction with large language models (LLMs). While LLM routing is effective in single-turn settings, ex

Generating Effective CoT Traces for Mitigating Causal Hallucination

TutorialsDGX agent

arXiv:2604.12748v1 Announce Type: new Abstract: Although large language models (LLMs) excel in complex reasoning tasks, they suffer from severe causal hallucination in event causality identification (

GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning

SafetyDGX agent

arXiv:2604.12630v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have exhibited remarkable performance in various visual tasks, yet still struggle with spatial reasoning. Rec

GigaCheck: Detecting LLM-generated Content via Object-Centric Span Localization

Local AiDGX agent

arXiv:2410.23728v3 Announce Type: replace Abstract: With the increasing quality and spread of LLM assistants, the amount of generated content is growing rapidly. In many cases and tasks, such texts ar

GLeMM: A large-scale multilingual dataset for morphological research

ResearchDGX agent

arXiv:2604.12442v1 Announce Type: new Abstract: In derivational morphology, what mechanisms govern the variation in form-meaning relations between words? The answers to this type of questions are typi

GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts

Model ReleasesDGX agent

arXiv:2604.12978v1 Announce Type: new Abstract: Optical character recognition (OCR) has advanced rapidly with the rise of vision-language models, yet evaluation has remained concentrated on a small cl

GRADE: Probing Knowledge Gaps in LLMs through Gradient Subspace Dynamics

ApplicationsDGX agent

arXiv:2604.02830v2 Announce Type: replace Abstract: Detecting whether a model's internal knowledge is sufficient to correctly answer a given question is a fundamental challenge in deploying responsibl

Gradient boundaries through confidence intervals for forced alignment estimates using model ensembles

SafetyDGX agent

arXiv:2506.01256v4 Announce Type: replace-cross Abstract: Forced alignment is a common tool to align audio with orthographic and phonetic transcriptions. Most forced alignment tools provide only point

Graph-Based Chain-of-Thought Pruning for Reducing Redundant Reflections in Reasoning LLMs

SafetyDGX agent

arXiv:2604.05643v2 Announce Type: replace Abstract: Extending CoT through RL has been widely used to enhance the reasoning capabilities of LLMs. However, due to the sparsity of reward signals, it can

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration

Model ReleasesDGX agent

arXiv:2604.12843v1 Announce Type: new Abstract: The rapid release of both language models and benchmarks makes it increasingly costly to evaluate every model on every dataset. In practice, models are

Hear Both Sides: Efficient Multi-Agent Debate via Diversity-Aware Message Retention

AgentsDGX agent

arXiv:2603.20640v2 Announce Type: replace Abstract: Multi-Agent Debate has emerged as a promising framework for improving the reasoning quality of large language models through iterative inter-agent c

InsightFlow: LLM-Driven Synthesis of Patient Narratives for Mental Health into Causal Models

SafetyDGX agent

arXiv:2604.12721v1 Announce Type: new Abstract: Clinical case formulation organizes patient symptoms and psychosocial factors into causal models, often using the 5P framework. However, constructing su

KCLarity at SemEval-2026 Task 6: Encoder and Zero-Shot Approaches to Political Evasion Detection

Model ReleasesDGX agent

arXiv:2603.06552v2 Announce Type: replace Abstract: This paper describes the KCLarity team's participation in CLARITY, a shared task at SemEval 2026 on classifying ambiguity and evasion techniques in

Knowledge Is Not Static: Order-Aware Hypergraph RAG for Language Models

ApplicationsDGX agent

arXiv:2604.12185v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enhances large language models by grounding outputs in retrieved knowledge. However, existing RAG methods including

KoCo: Conditioning Language Model Pre-training on Knowledge Coordinates

TutorialsDGX agent

arXiv:2604.12397v1 Announce Type: new Abstract: Standard Large Language Model (LLM) pre-training typically treats corpora as flattened token sequences, often overlooking the real-world context that hu

Latent-Condensed Transformer for Efficient Long Context Modeling

ResearchDGX agent

arXiv:2604.12452v1 Announce Type: new Abstract: Large language models (LLMs) face significant challenges in processing long contexts due to the linear growth of the key-value (KV) cache and quadratic

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Local AiDGX agent

arXiv:2601.14004v4 Announce Type: replace Abstract: Mechanistic Interpretability (MI) has emerged as a vital approach to demystify the opaque decision-making of Large Language Models (LLMs). However,

LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models

ResearchDGX agent

arXiv:2604.12056v1 Announce Type: new Abstract: Block-wise diffusion language models (DLMs) generate multiple tokens in any order, offering a promising alternative to the autoregressive decoding pipel

Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness

Local AiDGX agent

arXiv:2604.12373v1 Announce Type: new Abstract: Humans use introspection to evaluate their understanding through private internal states inaccessible to external observers. We investigate whether larg

Meet Dynamic Individual Preferences: Resolving Conflicting Human Value with Paired Fine-Tuning

SafetyDGX agent

arXiv:2604.12479v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have significantly improved the alignment of models with general human preferences. However, a major cha

MetFuse: Figurative Fusion between Metonymy and Metaphor

ResearchDGX agent

arXiv:2604.12919v1 Announce Type: new Abstract: Metonymy and metaphor often co-occur in natural language, yet computational work has studied them largely in isolation. We introduce a framework that tr

MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models

Model ReleasesDGX agent

arXiv:2604.12928v1 Announce Type: new Abstract: Speech-to-speech language models have recently emerged to enhance the naturalness of conversational AI. In particular, full-duplex models are distinguis

Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data

ResearchDGX agent

arXiv:2604.12633v1 Announce Type: new Abstract: Emotion classification in multilingual settings remains constrained by the scarcity of annotated data: existing corpora are predominantly English, singl

NaviRAG: Towards Active Knowledge Navigation for Retrieval-Augmented Generation

AgentsDGX agent

arXiv:2604.12766v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) typically relies on a flat retrieval paradigm that maps queries directly to static, isolated text segments. This ap

OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning

AgentsDGX agent

arXiv:2502.11271v2 Announce Type: replace-cross Abstract: Solving complex reasoning tasks may involve visual understanding, domain knowledge retrieval, numerical calculation, and multi-step reasoning.

Olmo 3

Model ReleasesDGX agent

arXiv:2512.13961v2 Announce Type: replace Abstract: We introduce Olmo 3, a family of state-of-the-art, fully-open language models at the 7B and 32B parameter scales. Olmo 3 model construction targets

ParetoBandit: Budget-Paced Adaptive Routing for Non-Stationary LLM Serving

Model ReleasesDGX agent

arXiv:2604.00136v2 Announce Type: replace-cross Abstract: Multi-model LLM serving operates in a non-stationary, noisy environment: providers revise pricing, model quality can shift or regress without

Perception-Aware Policy Optimization for Multimodal Reasoning

SafetyDGX agent

arXiv:2507.06448v5 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has proven to be a highly effective strategy for endowing Large Language Models (LLMs) with ro

PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models

ResearchDGX agent

arXiv:2601.19917v2 Announce Type: replace Abstract: Strategic planning is critical for multi-step reasoning, yet compact Large Language Models (LLMs) often lack the capacity to formulate global strate

PolicyLLM: Towards Excellent Comprehension of Public Policy for Large Language Models

Model ReleasesDGX agent

arXiv:2604.12995v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly integrated into real-world decision-making, including in the domain of public policy. Yet, their ability t

ReasonXL: Shifting LLM Reasoning Language Without Sacrificing Performance

Model ReleasesDGX agent

arXiv:2604.12378v1 Announce Type: new Abstract: Despite advances in multilingual capabilities, most large language models (LLMs) remain English-centric in their training and, crucially, in their produ

Representing expertise accelerates learning from pedagogical interaction data

ResearchDGX agent

arXiv:2604.12195v1 Announce Type: new Abstract: Work in cognitive science and artificial intelligence has suggested that exposing learning agents to traces of interaction between multiple individuals

← Previous
1…118119120121122…128
Next →