AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
21 Apr 2026

Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs

Model ReleasesDGX agent

arXiv:2604.18203v1 Announce Type: new Abstract: Multimodal LLMs can accurately perceive numerical content across modalities yet fail to perform exact multi-digit multiplication when the identical unde

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

ResearchDGX agent

arXiv:2505.20715v2 Announce Type: replace-cross Abstract: Video temporal understanding is crucial for multimodal large language models (MLLMs) to reason over events in videos. Despite recent advances

Navigating the Conceptual Multiverse

SafetyDGX agent

arXiv:2604.17815v1 Announce Type: cross Abstract: When language models answer open-ended problems, they implicitly make hidden decisions that shape their outputs, leaving users with uncontextualized a


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Negative Advantage Is a Double-Edged Sword: Calibrating Advantage in GRPO for Deep Search

SafetyDGX agent

arXiv:2604.18235v1 Announce Type: new Abstract: Deep search agents can autonomously initiate multi-turn interactions with search engines, thereby exhibiting strong question-answering capabilities. Suc

Neuro-Symbolic Resolution of Recommendation Conflicts in Multimorbidity Clinical Guidelines

Model ReleasesDGX agent

arXiv:2604.17340v1 Announce Type: new Abstract: Clinical guidelines, typically developed by independent specialty societies, inherently exhibit substantial fragmentation, redundancy, and logical contr

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR

Model ReleasesDGX agent

arXiv:2604.18105v1 Announce Type: cross Abstract: Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a mainstream paradigm in recent years. Although existing L

NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions

Model ReleasesDGX agent

arXiv:2604.16493v1 Announce Type: cross Abstract: Natural Language to SQL (NL2SQL) technology empowers non-expert users to query relational databases without requiring SQL expertise. While large langu

No One Fits All: From Fixed Prompting to Learned Routing in Multilingual LLMs

ResearchDGX agent

arXiv:2604.16937v1 Announce Type: new Abstract: Translation-based prompting is widely used in multilingual LLMs, yet its effectiveness varies across languages and tasks. We evaluate prompting strategi

No-Worse Context-Aware Decoding: Preventing Neutral Regression in Context-Conditioned Generation

ResearchDGX agent

arXiv:2604.16686v1 Announce Type: new Abstract: Large language models (LLMs) can answer questions and summarize documents when conditioned on external contexts (e.g., retrieved evidence), yet context

Oblivion: Self-Adaptive Agentic Memory Control through Decay-Driven Activation

AgentsDGX agent

arXiv:2604.00131v2 Announce Type: replace Abstract: Human memory adapts through selective forgetting: experiences become less accessible over time but can be reactivated by reinforcement or contextual

Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval

ApplicationsDGX agent

arXiv:2604.18360v1 Announce Type: cross Abstract: Audio-text retrieval systems based on Contrastive Language-Audio Pretraining (CLAP) achieve strong performance on traditional benchmarks; however, the

On Safety Risks in Experience-Driven Self-Evolving Agents

SafetyDGX agent

arXiv:2604.16968v1 Announce Type: new Abstract: Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self

On the Emergence of Syntax by Means of Local Interaction

Model ReleasesDGX agent

arXiv:2604.17857v1 Announce Type: new Abstract: Can syntactic processing emerge spontaneously from purely local interaction? We present a concrete instance on a minimal system: an 18,658-parameter two

On the Importance and Evaluation of Narrativity in Natural Language AI Explanations

Model ReleasesDGX agent

arXiv:2604.18311v1 Announce Type: new Abstract: Explainable AI (XAI) aims to make the behaviour of machine learning models interpretable, yet many explanation methods remain difficult to understand. T

On the Predictive Power of Representation Dispersion in Language Models

Model ReleasesDGX agent

arXiv:2506.24106v2 Announce Type: replace Abstract: We show that a language model's ability to predict text is tightly linked to the breadth of its embedding space: models that spread their contextual

On the Robustness of LLM-Based Dense Retrievers: A Systematic Analysis of Generalizability and Stability

ResearchDGX agent

arXiv:2604.16576v1 Announce Type: cross Abstract: Decoder-only large language models (LLMs) are increasingly replacing BERT-style architectures as the backbone for dense retrieval, achieving substanti

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization

SafetyDGX agent

arXiv:2509.23542v2 Announce Type: replace Abstract: The LLM-as-a-judge paradigm is widely used in both evaluating free-text model responses and reward modeling for model alignment and fine-tuning. Rec

One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment

SafetyDGX agent

arXiv:2601.18731v2 Announce Type: replace Abstract: Alignment of Large Language Models (LLMs) aims to align outputs with human preferences, and personalized alignment further adapts models to individu

OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

AgentsDGX agent

arXiv:2604.18486v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) reasoning has become a powerful driver of trajectory prediction in VLA-based autonomous driving, yet its autoregressive nature

ONTO: A Token-Efficient Columnar Notation for LLM Input Optimization

ResearchDGX agent

arXiv:2604.17512v1 Announce Type: new Abstract: Serialization formats designed for document interchange impose structural overhead that becomes prohibitive when large language models consume operation

OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation

Model ReleasesDGX agent

arXiv:2506.05606v5 Announce Type: replace Abstract: Can large language models (LLMs) accurately simulate the next web action of a specific user? While LLMs have shown promising capabilities in generat

OPSDL: On-Policy Self-Distillation for Long-Context Language Models

SafetyDGX agent

arXiv:2604.17535v1 Announce Type: new Abstract: Extending the effective context length of large language models (LLMs) remains a central challenge for real-world applications. While recent post-traini

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist

ResearchDGX agent

arXiv:2601.02986v2 Announce Type: replace Abstract: Recent approaches in personalized reward modeling have primarily focused on leveraging user interaction history to align model judgments with indivi

Parallel Test-Time Scaling for Latent Reasoning Models

Model ReleasesDGX agent

arXiv:2510.07745v4 Announce Type: replace Abstract: Parallel test-time scaling (TTS) is a pivotal approach for enhancing large language models (LLMs), typically by sampling multiple token-based chains

PARM: Pipeline-Adapted Reward Model

ApplicationsDGX agent

arXiv:2604.18327v1 Announce Type: cross Abstract: Reward models (RMs) are central to aligning large language models (LLMs) with human preferences, powering RLHF and advanced decoding strategies. While

PDDL-Mind: Large Language Models are Capable on Belief Reasoning with Reliable State Tracking

Model ReleasesDGX agent

arXiv:2604.17819v1 Announce Type: new Abstract: Large language models (LLMs) perform substantially below human level on existing theory-of-mind (ToM) benchmarks, even when augmented with chain-of-thou

Pearmut: Human Evaluation of Translation Made Trivial

ResearchDGX agent

arXiv:2601.02933v3 Announce Type: replace Abstract: Human evaluation is the gold standard for multilingual NLP, but is often skipped in practice and substituted with automatic metrics because it is no

Peerispect: Claim Verification in Scientific Peer Reviews

SafetyDGX agent

arXiv:2604.17667v1 Announce Type: new Abstract: Peer review is central to scientific publishing, yet reviewers frequently include claims that are subjective, rhetorical, or misaligned with the submitt

PersonalHomeBench: Evaluating Agents in Personalized Smart Homes

Model ReleasesDGX agent

arXiv:2604.16813v1 Announce Type: cross Abstract: Agentic AI systems are rapidly advancing toward real-world applications, yet their readiness in complex and personalized environments remains insuffic

Personalizing Student-Agent Interactions Using Log-Contextualized Retrieval-Augmented Generation (RAG)

AgentsDGX agent

arXiv:2505.17238v3 Announce Type: replace Abstract: Collaborative dialogue offers rich insights into students' learning and critical thinking, which is essential for personalizing pedagogical agent in

PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning

HardwareDGX agent

arXiv:2509.18169v3 Announce Type: replace-cross Abstract: Tasks on complex systems require high-precision numerical computation to support decisions, but current large language models (LLMs) cannot in

Plausibility as Commonsense Reasoning: Humans Succeed, Large Language Models Do not

TutorialsDGX agent

arXiv:2604.04825v2 Announce Type: replace Abstract: Large language models achieve strong performance on many language tasks, yet it remains unclear whether they integrate world knowledge with syntacti

Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive Decoding

Model ReleasesDGX agent

arXiv:2604.17132v1 Announce Type: new Abstract: Safety-aligned large language models (LLMs) often generate refusal responses to harmless queries due to the over-refusal problem. However, existing meth

PoliLegalLM: A Technical Report on a Large Language Model for Political and Legal Affairs

SafetyDGX agent

arXiv:2604.17543v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable success in general-domain tasks, yet their direct application to the legal domain remains challeng

Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs

Model ReleasesDGX agent

arXiv:2604.17837v1 Announce Type: cross Abstract: An LLM's residual stream is both state and instruction: it encodes the current context and determines the next transformation. We introduce a paramete

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning

ResearchDGX agent

arXiv:2502.02871v2 Announce Type: replace Abstract: Scientific reasoning, the process through which humans apply logic, evidence, and critical thinking to explore and interpret scientific phenomena, i

Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?

Model ReleasesDGX agent

arXiv:2604.17338v1 Announce Type: cross Abstract: Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but o

PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention

Model ReleasesDGX agent

arXiv:2506.13674v3 Announce Type: replace Abstract: Parameter-Efficient Fine-Tuning (PEFT) methods have become crucial for rapidly adapting large language models (LLMs) to downstream tasks. Prefix-Tun

PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment

SafetyDGX agent

arXiv:2508.05132v2 Announce Type: replace Abstract: As medical LLMs transition to clinical deployment, assessing their ethical reasoning capability becomes critical. While achieving high accuracy on k

PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations

Model ReleasesDGX agent

arXiv:2604.16909v1 Announce Type: new Abstract: As large language models (LLMs) evolve from conversational assistants into agents capable of handling complex tasks, they are increasingly deployed in h

PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues

ResearchDGX agent

arXiv:2604.18354v1 Announce Type: new Abstract: Emotion plays a pivotal role in shaping negotiation outcomes, influencing trust, cooperation, and long-term relationships. Developing negotiation dialog

Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models

SafetyDGX agent

arXiv:2601.15220v2 Announce Type: replace Abstract: We identify a novel phenomenon in language models: benign fine-tuning of frontier models can lead to privacy collapse. We find that diverse, subtle

Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning

Local AiDGX agent

arXiv:2510.16054v2 Announce Type: replace-cross Abstract: When users submit queries to Large Language Models (LLMs), their prompts can often contain sensitive data, forcing a difficult choice: Send th

PRL: Prompts from Reinforcement Learning

ResearchDGX agent

arXiv:2505.14412v2 Announce Type: replace-cross Abstract: Effective prompt engineering remains a central challenge in fully harnessing the capabilities of LLMs. While well-designed prompts can dramati

Probabilistic Programs of Thought

HardwareDGX agent

arXiv:2604.17290v1 Announce Type: new Abstract: LLMs are widely used for code generation and mathematical reasoning tasks where they are required to generate structured output. They either need to rea

Procedural Knowledge at Scale Improves Reasoning

TutorialsDGX agent

arXiv:2604.01348v2 Announce Type: replace Abstract: Test-time scaling has emerged as an effective way to improve language models on challenging reasoning tasks. However, most existing methods treat ea

Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards

ResearchDGX agent

arXiv:2604.17957v1 Announce Type: new Abstract: Process Reward Models (PRMs) have emerged as a powerful tool for providing step-level feedback when evaluating the reasoning of Large Language Models (L

ProfVLM: A lightweight video-language model for multi-view proficiency estimation

Model ReleasesDGX agent

arXiv:2509.26278v4 Announce Type: replace-cross Abstract: Most existing approaches formulate action quality assessment and skill proficiency estimation as discriminative prediction tasks, typically pr

Prune, Interpret, Evaluate: A Cross-Layer Transcoder-Native Framework for Efficient Circuit Discovery via Feature Attribution

Model ReleasesDGX agent

arXiv:2604.16889v1 Announce Type: new Abstract: Existing feature-interpretation pipelines typically operate on uniformly sampled units, but only a small fraction of cross-layer transcoder (CLT) featur

Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition

Model ReleasesDGX agent

arXiv:2510.08047v2 Announce Type: replace-cross Abstract: Robust ASR under domain shift is crucial because real-world systems encounter unseen accents and domains with limited labeled data. Although p

QU-NLP at QIAS 2026: Multi-Stage QLoRA Fine-Tuning for Arabic Islamic Inheritance Reasoning

Model ReleasesDGX agent

arXiv:2604.16396v1 Announce Type: new Abstract: Islamic inheritance law (ilm al-mawar{i}th) presents a challenging domain for evaluating large language models' structured reasoning capabilities, requi

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks

ResearchDGX agent

arXiv:2604.17842v1 Announce Type: new Abstract: LLM benchmarks are increasingly dynamic: instead of containing a fixed set of questions, they define templates and parameters that can generate an effec

RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction

ApplicationsDGX agent

arXiv:2504.07415v2 Announce Type: replace-cross Abstract: Automated radiology report generation (RRG) holds potential to reduce the workload of radiologists, and recent advances in multimodal large la

RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation

ApplicationsDGX agent

arXiv:2604.16310v1 Announce Type: cross Abstract: Evaluating Retrieval-Augmented Generation (RAG) systems using static multi-turn datasets fails to capture the dynamic nature of real-world dialogues.

ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval

Model ReleasesDGX agent

arXiv:2510.08252v2 Announce Type: replace-cross Abstract: In this paper, we introduce ReasonEmbed, a novel text embedding model developed for reasoning-intensive document retrieval. Our work includes

Reasoning Models Know What's Important, and Encode It in Their Activations

ResearchDGX agent

arXiv:2604.18307v1 Announce Type: new Abstract: Language models often solve complex tasks by generating long reasoning chains, consisting of many steps with varying importance. While some steps are cr

Reciprocal Co-Training (RCT): Coupling Gradient-Based and Non-Differentiable Models via Reinforcement Learning

TutorialsDGX agent

arXiv:2604.16378v1 Announce Type: new Abstract: Large language models (LLMs) and classical machine learning methods offer complementary strengths for predictive modeling, yet their fundamentally diffe

ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering

Model ReleasesDGX agent

arXiv:2604.17944v1 Announce Type: new Abstract: Developing agents capable of navigating fragmented, multi-source information remains challenging, primarily due to the scarcity of benchmarks reflecting

REFLEX: Reference-Free Evaluation of Log Summarization via Large Language Model Judgment

ApplicationsDGX agent

arXiv:2511.07458v2 Announce Type: replace Abstract: Evaluating log summarization systems is challenging due to the lack of high-quality reference summaries and the limitations of existing metrics like

REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control

Model ReleasesDGX agent

arXiv:2511.20233v3 Announce Type: replace Abstract: The prevalence of fake news on social media demands automated fact-checking systems to provide accurate verdicts with faithful explanations. However

← Previous
1…109110111112113…129
Next →