AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
26 May 2026

ECHO: Terminal Agents Learn World Models for Free

SafetyDGX agent

arXiv:2605.24517v1 Announce Type: cross Abstract: CLI agents are the closest thing language models have to an embodied setting: the model emits commands, the terminal executes them, and the returned s

EfficientGraph-RAG: Structured Retrieval-State Management for Cross-Task Retrieval-Augmented Generation

AgentsDGX agent

arXiv:2605.25379v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) has become the standard way to ground large language models in external knowledge, but many systems still organize

End-to-End Intracortical Speech Decoding from Neural Activity

ResearchDGX agent

arXiv:2605.24313v1 Announce Type: new Abstract: Current high-performing intracortical speech neuroprostheses achieve low word error rates but typically rely on external language models during inferenc


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Exploring Profiles of Cognitive Distortions Associated with Mental Health Disorders

ResearchDGX agent

arXiv:2605.24996v1 Announce Type: new Abstract: Cognitive distortions, distorted patterns of thinking, have been increasingly studied in computational mental health research. Although they are related

Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges

SafetyDGX agent

arXiv:2605.23970v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as automatic judges for summarization and dialogue evaluation. Prior work has documented biases such

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning

SafetyDGX agent

arXiv:2605.24286v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning is useful for monitoring language models only when the reasoning trace faithfully reflects the computation that produ

Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth

Model ReleasesDGX agent

arXiv:2605.25052v1 Announce Type: new Abstract: Chains of thought (CoTs) have become central in interpreting and auditing behaviors of large language models. Yet growing evidence suggests that these t

Feature Resemblance: Towards a Theoretical Understanding of Analogical Reasoning in Transformers

ResearchDGX agent

arXiv:2603.05143v3 Announce Type: replace Abstract: Understanding reasoning in large language models is complicated by evaluations that conflate multiple reasoning types. We isolate analogical reasoni

Forgotten Words: Benchmarking NeoBERT for Dementia Detection in Low-Resource Conversational Filipino and English Speech

ResearchDGX agent

arXiv:2605.26007v1 Announce Type: new Abstract: Dementia detection from spontaneous speech offers a scalable approach to cognitive screening, yet NLP systems remain predominantly English-centric. This

Found in Conversation: LLMs Teach Themselves to Close the Multi-Turn Gap

Model ReleasesDGX agent

arXiv:2605.24432v1 Announce Type: new Abstract: Large Language Model (LLM) interactions are typically underspecified, with users clarifying all necessary details across multiple conversational turns.

From Automation to Collaboration: Human-in-the-Loop Methods for Safe and Trustworthy NLP

SafetyDGX agent

arXiv:2605.25226v1 Announce Type: new Abstract: Large language models are widely deployed in high-stakes NLP tasks, yet risks such as bias, hallucination, adversarial vulnerability and unreliable gene

From Facts to Insights: A Persona-Driven Dual Memory Framework and Dataset for Role-Playing Agents

Model ReleasesDGX agent

arXiv:2605.25693v1 Announce Type: new Abstract: While role-playing agents excel in short-term interactions, long-term conversations overwhelm context windows, motivating external memory frameworks. Cu

From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlas

SafetyDGX agent

arXiv:2602.00491v2 Announce Type: replace Abstract: Public health reasoning requires population level inference grounded in scientific evidence, expert consensus, and safety constraints. However, it r

Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation

ApplicationsDGX agent

arXiv:2605.24534v1 Announce Type: new Abstract: We present a fully automated pipeline that transforms large collections of court decisions into legal commentaries for statutes - without providing any

GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving

ResearchDGX agent

arXiv:2605.25384v1 Announce Type: new Abstract: Mathematical reasoning is a hallmark of human intelligence, requiring logical deduction, symbolic manipulation, and abstract thinking. Recent multimodal

GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation

SafetyDGX agent

arXiv:2605.25447v1 Announce Type: new Abstract: Generating structured, editable diagrams remains a significant challenge for contemporary large language models, despite their proficiency in general-pu

GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning

Model ReleasesDGX agent

arXiv:2605.25200v1 Announce Type: new Abstract: Travel planning is a realistic task for evaluating the planning and tool-use abilities of LLM agents. However, existing benchmarks typically assume only

H^{2}MT: Semantic Hierarchy-Aware Hierarchical Memory Transformer

HardwareDGX agent

arXiv:2605.24930v1 Announce Type: new Abstract: Transformer-based LLMs achieve strong results on many language tasks; however, long inputs remain challenging because context windows are finite, and pr

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models

SafetyDGX agent

arXiv:2605.25443v1 Announce Type: new Abstract: Post-training has significantly enhanced the reasoning capability of Large Reasoning Models (LRMs), especially with Reinforcement Learning (RL) like Gro

Hierarchical Local-Global Transformer for Temporal Sentence Grounding

Local AiDGX agent

arXiv:2208.14882v2 Announce Type: replace-cross Abstract: This paper studies the multimedia problem of temporal sentence grounding (TSG), which aims to accurately determine the specific video segment

HiMed: Incentivizing Hindi Reasoning in Medical LLMs

Model ReleasesDGX agent

arXiv:2605.24635v1 Announce Type: new Abstract: Medical large language models hold promise for reducing healthcare disparities, yet Hindi remains severely underrepresented. While medical LLMs excel in

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework

Model ReleasesDGX agent

arXiv:2507.19219v2 Announce Type: replace Abstract: Overestimation in evaluating large language models (LLMs) has become an increasing concern. Due to the contamination of public benchmarks or imbalan

How Much Structure Do LLMs Need? Evaluating LLMs for Bibliometric Cluster Description

SafetyDGX agent

arXiv:2605.24351v1 Announce Type: new Abstract: Large language models (LLMs) can support scientific literature synthesis, but remain prone to hallucinated references, uneven coverage, and weakly groun

HyLaT: Efficient Multi-Agent Communication via Hybrid Latent-Text Protocol

AgentsDGX agent

arXiv:2605.25421v1 Announce Type: new Abstract: Communication protocol design is a central challenge in large language model-based multi-agent systems. Existing single-channel approaches face an inher

Improving the Completeness and Comparability of Segment Disclosures: A Large Language Model Approach

SafetyDGX agent

arXiv:2605.23924v1 Announce Type: new Abstract: Segment-level disclosures are a central component of financial reporting, providing insight into firms' internal organization and the allocation of econ

Ineffectiveness for Search and Undecidability of PCSP Meta-Problems

ResearchDGX agent

arXiv:2504.04639v4 Announce Type: replace-cross Abstract: It is an open question whether the search and decision versions of promise CSPs are equivalent. Most known algorithms for PCSPs solve only the

Inference Time Optimization with Confidence Dynamics

Model ReleasesDGX agent

arXiv:2605.25244v1 Announce Type: new Abstract: Inference time optimization techniques, such as repeated sampling, have significantly advanced the reasoning capabilities of Large Language Models (LLMs

InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models

Model ReleasesDGX agent

arXiv:2505.13878v3 Announce Type: replace-cross Abstract: Model fusion combines multiple Large Language Models (LLMs) with different strengths into a more powerful, integrated model through lightweigh

Is Inference Mediated by Distinct Semantic Structures in LLMs? A Mechanistic Interpretation

ResearchDGX agent

arXiv:2605.25520v1 Announce Type: new Abstract: Predicting a label correctly does not necessarily require representing the operation that produces it. Transformer representations are known to carry la

Iterate Until Retrieved: Factual Nugget Optimization for Discoverable Continual Corrections in Agentic RAG

AgentsDGX agent

arXiv:2605.25641v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) systems in complex B2B (business-to-business) settings may often receive free-form response feedback. Rathe

Know You Before You Speak: User-State Modeling for LLM Personalization in Multi-Turn Conversation

ApplicationsDGX agent

arXiv:2605.24647v1 Announce Type: new Abstract: Personalized dialogue requires more than recalling explicit user histories: systems also need to infer hidden user states that evolve through interactio

Knowing but Not Showing: LLMs Recognize Ambiguity but Rarely Ask Clarifying Questions

ResearchDGX agent

arXiv:2605.25284v1 Announce Type: new Abstract: User queries are often underspecified and may admit multiple valid interpretations. Rather than silently making assumptions about the user's intent, a h

Large Language Model Selection with Limited Annotations

ResearchDGX agent

arXiv:2605.24981v1 Announce Type: new Abstract: Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotations o

Lean Formalization of Generalization Error Bound by Rademacher Complexity and Dudley's Entropy Integral

ResearchDGX agent

arXiv:2503.19605v5 Announce Type: replace-cross Abstract: Understanding and certifying the generalization performance of machine learning algorithms -- i.e. obtaining theoretical estimates of the test

Learning Concepts, Not Tokens: Self-Supervised Semantic Alignment for Language Models

SafetyDGX agent

arXiv:2603.29123v2 Announce Type: replace Abstract: The next-token prediction (NTP) objective trains language models to predict a single token at each step, even though many continuations can express

Learning to Route Languages for Multilingual Policy Optimization

SafetyDGX agent

arXiv:2605.25360v1 Announce Type: new Abstract: Large language models~(LLMs) are trained on heterogeneous multilingual corpora, yet existing policy optimization methods often implicitly restrict each

Llamion Technical Report

Model ReleasesDGX agent

arXiv:2605.25676v1 Announce Type: new Abstract: We release Llamion, a family of 14B-parameter open-weight language models obtained by transforming Orion-14B into the standardized Llama-family architec

LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers

Model ReleasesDGX agent

arXiv:2605.25415v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in academic peer review, yet their reliability, alignment with human judgment, and robustness to adve

Lngram: N-gram Conditional Memory in Latent Space

Local AiDGX agent

arXiv:2605.24869v1 Announce Type: new Abstract: Sequence modeling requires both compositional reasoning and local static knowledge retrieval, yet standard Transformers handle both through dense comput

Locality Matters for Training-Free Audio Token Compression in Audio-Language Models

SafetyDGX agent

arXiv:2605.25179v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used for audio captioning, question answering, and open-ended audio understanding, but their inference cos

MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models

SafetyDGX agent

arXiv:2605.26004v1 Announce Type: cross Abstract: Instruction tuning of large vision-language models (LVLMs) increasingly depends on massive multimodal corpora, yet these datasets contain samples with

Mapping the Schedule x Bit-Width Boundary in Sub-100M Quantisation-Aware Training

ResearchDGX agent

arXiv:2605.25966v1 Announce Type: cross Abstract: We test whether the optimal learning-rate schedule depends on bit-width during from-initialisation quantisation-aware training (QAT) for sub-100M deco

MATO: Multi-objective Personalized Alignment with Test-time Optimization for Large Language Models

SafetyDGX agent

arXiv:2605.25342v1 Announce Type: new Abstract: Aligning large language models (LLMs) with diverse and multifaceted user preferences is a fundamental challenge in personalized AI systems. Existing mul

MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding

Model ReleasesDGX agent

arXiv:2605.24523v1 Announce Type: cross Abstract: Visual decoding from brain signals is a key challenge at the intersection of computer vision and neuroscience, requiring methods that bridge neural re

Mitigating Hallucinations in Healthcare LLMs with Granular Fact-Checking and Domain-Specific Adaptation

SafetyDGX agent

arXiv:2512.16189v3 Announce Type: replace Abstract: In healthcare, it is essential for any LLM-generated output to be reliable and accurate, particularly in cases involving decision-making and patient

Mitigating Provenance-Role Collapse in Long-Term Agents via Typed Memory Representation

ResearchDGX agent

arXiv:2605.25869v1 Announce Type: new Abstract: Long-term memory is essential for persistent LLM agents, yet prevailing architectures store historical interactions as unstructured, flat text. This unc

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Model ReleasesDGX agent

arXiv:2505.23764v3 Announce Type: replace-cross Abstract: Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, h

Multi-Persona Debate System for Automated Scientific Hypothesis Generation

AgentsDGX agent

arXiv:2605.23917v1 Announce Type: new Abstract: Modern scientific discovery is bottlenecked not by data scarcity, but by the inability to synthesize fragmented knowledge into actionable hypotheses. Th

MultiHaluDet: Multilingual Hallucination Detection via LLM Hidden State Probing

Model ReleasesDGX agent

arXiv:2605.24919v1 Announce Type: new Abstract: Hallucinations in Large Language Models (LLMs) represent a critical barrier to their reliable deployment, a vulnerability heavily exacerbated in non-Eng

Multilingual Phonological Feature Recognition with Self-Supervised Speech Models

ResearchDGX agent

arXiv:2605.25596v1 Announce Type: new Abstract: Phonological features provide a language-general and linguistically grounded representation of speech. We present PhonoQ-2.0, a multilingual frame-level

Neural Router: Semantic Content Matching for Agentic AI

Model ReleasesDGX agent

arXiv:2605.25701v1 Announce Type: cross Abstract: Large language models (LLMs) can serve as the semantic-matching engine of a content-based publish/subscribe broker for agentic AI across the edge-clou

NITP: Next Implicit Token Prediction for LLM Pre-training

ResearchDGX agent

arXiv:2605.24956v1 Announce Type: new Abstract: Standard next-token prediction (NTP) supervises language models solely through discrete labels in the output logit space. We argue that this sparse one-

On the Limits of Model Merging for Multilinguality in Pre-Training

ResearchDGX agent

arXiv:2605.25846v1 Announce Type: new Abstract: Endowing models with consistent multilingual performance can be achieved by mixing pre-training data, or post-training approaches such as language-speci

Optimizing Token Choice for Code Watermarking: An RL Approach

SafetyDGX agent

arXiv:2508.11925v3 Announce Type: replace-cross Abstract: Protecting intellectual property on LLM-generated code necessitates effective watermarking systems that can operate within code's highly struc

Overview of the PsyDefDetect Shared Task at BioNLP 2026: Detecting Levels of Psychological Defense Mechanisms in Supportive Conversations

Model ReleasesDGX agent

arXiv:2605.24907v1 Announce Type: new Abstract: We present an overview of PsyDefDetect, the shared task on detecting levels of psychological defense mechanisms in emotional support dialogues, co-locat

P1SCO: Social Dimensions from a Perspectivist Lens

ResearchDGX agent

arXiv:2605.25312v1 Announce Type: new Abstract: We introduce P1SCO, a dataset of social media comments collected from three distinct platforms, annotated according to ten social dimensions to capture

Peak-Then-Collapse and the Four Interface Channels of Knowledge-Graph Tool Use

SafetyDGX agent

arXiv:2605.26037v1 Announce Type: new Abstract: We test the standard RLVR tool-use recipe -- GRPO on Qwen2.5-7B-Instruct -- on a deliberately minimal knowledge-graph tool API: four Freebase navigation

PerSoMed: A Large-Scale Balanced Dataset for Persian Social Media Text Classification

ResearchDGX agent

arXiv:2602.19333v2 Announce Type: replace Abstract: This research introduces the first large-scale, well-balanced Persian social media text classification dataset, specifically designed to address the

Persuasion Should be Double-Blind: A Multi-Domain Dialogue Dataset With Faithfulness Based on Causal Theory of Mind

AgentsDGX agent

arXiv:2502.21297v2 Announce Type: replace Abstract: Persuasive dialogue is central to human communication, yet existing datasets often rely on a single language model generating both roles, producing

Phonetic Modeling of Dialectal Variation in Vietnamese Speech

ResearchDGX agent

arXiv:2605.24451v1 Announce Type: new Abstract: Vietnamese exhibits substantial dialectal phonetic variation across Northern, Central, and Southern regions, where identical lexical items may be realiz

← Previous
1…6061626364…129
Next →