AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
15 Apr 2026

Retrieval as a Decision: Training-Free Adaptive Gating for Efficient RAG

SafetyDGX agent

arXiv:2511.09803v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) improves factuality but retrieving for every query often hurts quality while inflating tokens and latency. We p

SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis

ResearchDGX agent

arXiv:2604.13035v1 Announce Type: cross Abstract: Large Language Models (LLMs) and Vision-Language Models (VLMs) increasingly generate indoor scenes through intermediate structures such as layouts and

Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

SafetyDGX agent

arXiv:2604.12002v1 Announce Type: new Abstract: Current post-training methods in verifiable settings fall into two categories. Reinforcement learning (RLVR) relies on binary rewards, which are broadly


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Sparse Growing Transformer: Training-Time Sparse Depth Allocation via Progressive Attention Looping

Model ReleasesDGX agent

arXiv:2603.23998v2 Announce Type: replace Abstract: Existing approaches to increasing the effective depth of Transformers predominantly rely on parameter reuse, extending computation through recursive

Speaker effects in language comprehension: An integrative model of language and speaker processing

ResearchDGX agent

arXiv:2412.07238v3 Announce Type: replace Abstract: The identity of a speaker influences language comprehension through modulating perception and expectation. This review explores speaker effects and

StoryScope: Investigating idiosyncrasies in AI fiction

Model ReleasesDGX agent

arXiv:2604.03136v4 Announce Type: replace Abstract: As AI-generated fiction becomes increasingly prevalent, questions of authorship and originality are becoming central to how written work is evaluate

Teaching LLMs Human-Like Editing of Inappropriate Argumentation via Reinforcement Learning

SafetyDGX agent

arXiv:2604.12770v1 Announce Type: new Abstract: Editing human-written text has become a standard use case of large language models (LLMs), for example, to make one's arguments more appropriate for a d

Temporal Flattening in LLM-Generated Text: Comparing Human and LLM Writing Trajectories

ResearchDGX agent

arXiv:2604.12097v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in daily applications, from content generation to code writing, where each interaction treats the mod

The Effect of Document Selection on Query-focused Text Analysis

ResearchDGX agent

arXiv:2604.12099v1 Announce Type: cross Abstract: Analyses of document collections often require selecting what data to analyze, as not all documents are relevant to a particular research question and

The Enforcement and Feasibility of Hate Speech Moderation on Twitter

ResearchDGX agent

arXiv:2604.12289v1 Announce Type: cross Abstract: Online hate speech is associated with substantial social harms, yet it remains unclear how consistently platforms enforce hate speech policies or whet

The role of System 1 and System 2 semantic memory structure in human and LLM biases

SafetyDGX agent

arXiv:2604.12816v1 Announce Type: new Abstract: Implicit biases in both humans and large language models (LLMs) pose significant societal risks. Dual process theories propose that biases arise primari

Think Through Uncertainty: Improving Long-Form Generation Factuality via Reasoning Calibration

ResearchDGX agent

arXiv:2604.12046v1 Announce Type: new Abstract: Large language models (LLMs) often hallucinate in long-form generation. Existing approaches mainly improve factuality through post-hoc revision or reinf

Thought-Retriever: Don't Just Retrieve Raw Data, Retrieve Thoughts for Memory-Augmented Agentic Systems

Model ReleasesDGX agent

arXiv:2604.12231v1 Announce Type: new Abstract: Large language models (LLMs) have transformed AI research thanks to their powerful internal capabilities and knowledge. However, existing LLMs still fai

TimeMark: A Trustworthy Time Watermarking Framework for Exact Generation-Time Recovery from AIGC

ResearchDGX agent

arXiv:2604.12216v1 Announce Type: cross Abstract: The widespread use of Large Language Models (LLMs) in text generation has raised increasing concerns about intellectual property disputes. Watermarkin

Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Sequence-Level Likelihood

SafetyDGX agent

arXiv:2604.12736v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has significantly advanced the reasoning ability of large language models (LLMs), particularly in their mathem

Toward Autonomous Long-Horizon Engineering for ML Research

AgentsDGX agent

arXiv:2604.13018v1 Announce Type: new Abstract: Autonomous AI research has advanced rapidly, but long-horizon ML research engineering remains difficult: agents must sustain coherent progress across ta

Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector

Model ReleasesDGX agent

arXiv:2509.07177v3 Announce Type: replace Abstract: Large language models have demonstrated impressive capabilities across various domains. However, their general-purpose nature often limits their eff

Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format Reasoning

AgentsDGX agent

arXiv:2604.12282v1 Announce Type: new Abstract: Spreadsheets are central to real-world applications such as enterprise reporting, auditing, and scientific data management. Despite their ubiquity, exis

ToxiTrace: Gradient-Aligned Training for Explainable Chinese Toxicity Detection

Model ReleasesDGX agent

arXiv:2604.12321v1 Announce Type: new Abstract: Existing Chinese toxic content detection methods mainly target sentence-level classification but often fail to provide readable and contiguous toxic evi

Transforming External Knowledge into Triplets for Enhanced Retrieval in RAG of LLMs

Model ReleasesDGX agent

arXiv:2604.12610v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) mitigates hallucination in large language models (LLMs) by incorporating external knowledge during generation. Howe

UCS: Estimating Unseen Coverage for Improved In-Context Learning

ResearchDGX agent

arXiv:2604.12015v1 Announce Type: cross Abstract: In-context learning (ICL) performance depends critically on which demonstrations are placed in the prompt, yet most existing selectors prioritize heur

Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark

Model ReleasesDGX agent

arXiv:2604.12744v1 Announce Type: new Abstract: While multilingual language models promise to bring the benefits of LLMs to speakers of many languages, gold-standard evaluation benchmarks in most lang

Using Learning Progressions to Guide AI Feedback for Science Learning

TutorialsDGX agent

arXiv:2603.03249v2 Announce Type: replace Abstract: Generative artificial intelligence (AI) offers scalable support for formative feedback, yet most AI-generated feedback relies on task-specific rubri

When Self-Reference Fails to Close: Matrix-Level Dynamics in Large Language Models

Model ReleasesDGX agent

arXiv:2604.12128v1 Announce Type: new Abstract: We investigate how self-referential inputs alter the internal matrix dynamics of large language models. Measuring 106 scalar metrics across up to 7 anal

WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering

Model ReleasesDGX agent

arXiv:2604.05818v2 Announce Type: replace-cross Abstract: Multi-modal Retrieval-Augmented Generation (RAG) has emerged as a highly effective paradigm for Knowledge-Based Visual Question Answering (KB-

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching

Model ReleasesDGX agent

arXiv:2507.09318v2 Announce Type: replace-cross Abstract: Generating spoken dialogue is inherently more complex than monologue text-to-speech (TTS), as it demands both realistic turn-taking and the ma

14 Apr 2026

A Multilingual Dataset and Empirical Validation for the Mutual Reinforcement Effect in Information Extraction

SafetyDGX agent

arXiv:2407.10953v5 Announce Type: replace Abstract: The Mutual Reinforcement Effect (MRE) describes a phenomenon in information extraction where word-level and sentence-level tasks can mutually improv

A Structured Clustering Approach for Inducing Media Narratives

ResearchDGX agent

arXiv:2604.10368v1 Announce Type: new Abstract: Media narratives wield tremendous power in shaping public opinion, yet computational approaches struggle to capture the nuanced storytelling structures

Adaptive Multi-Expert Reasoning via Difficulty-Aware Routing and Uncertainty-Guided Aggregation

ResearchDGX agent

arXiv:2604.10335v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate strong performance in math reasoning benchmarks, but their performance varies inconsistently across problems wi

Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning

Model ReleasesDGX agent

arXiv:2601.03641v4 Announce Type: replace Abstract: Large Language Model (LLM)-based agents significantly extend the utility of LLMs by interacting with dynamic environments. However, enabling agents

Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks

Model ReleasesDGX agent

arXiv:2604.11753v1 Announce Type: new Abstract: We study parallel test-time scaling for long-horizon agentic tasks such as agentic search and deep research, where multiple rollouts are generated in pa

Agri-R1: Agricultural Reasoning for Disease Diagnosis via Automated-Synthesis and Reinforcement Learning

Model ReleasesDGX agent

arXiv:2601.04672v2 Announce Type: replace-cross Abstract: Agricultural disease diagnosis challenges VLMs, as conventional fine-tuning requires extensive labels, lacks interpretability, and generalizes

Aligning What LLMs Do and Say: Towards Self-Consistent Explanations

Model ReleasesDGX agent

arXiv:2506.07523v3 Announce Type: replace Abstract: Large language models (LLMs) seem to offer an easy path to interpretability: just ask them to explain their answers. Yet the features driving an ans

Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models

ResearchDGX agent

arXiv:2604.10697v1 Announce Type: new Abstract: Large language models frequently exhibit hallucinations: fluent and confident outputs that are factually incorrect or unsupported by the input context.

AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption

Model ReleasesDGX agent

arXiv:2508.03793v2 Announce Type: replace Abstract: Long-context large language models (LLMs), such as Gemini-2.5-Pro and Claude-Sonnet-4, are increasingly used to empower advanced AI systems, includi

Back to Basics: Let Conversational Agents Remember with Just Retrieval and Generation

ResearchDGX agent

arXiv:2604.11628v1 Announce Type: new Abstract: Existing conversational memory systems rely on complex hierarchical summarization or reinforcement learning to manage long-term dialogue history, yet re

BadGraph: A Backdoor Attack Against Latent Diffusion Model for Text-Guided Graph Generation

Model ReleasesDGX agent

arXiv:2510.20792v4 Announce Type: replace-cross Abstract: The rapid progress of graph generation has raised new security concerns, particularly regarding backdoor vulnerabilities. While prior work has

Beyond Black-Box Interventions: Latent Probing for Faithful Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2510.12460v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) systems often fail to maintain contextual faithfulness, generating responses that conflict with the provided co

Beyond the Beep: Scalable Collision Anticipation and Real-Time Explainability with BADAS-2.0

Model ReleasesDGX agent

arXiv:2604.05767v2 Announce Type: replace-cross Abstract: We present BADAS-2.0, the second generation of our collision anticipation system, building on BADAS-1.0, which showed that fine-tuning V-JEPA2

BiT-MCTS: A Theme-based Bidirectional MCTS Approach to Chinese Fiction Generation

ResearchDGX agent

arXiv:2603.14410v3 Announce Type: replace Abstract: Generating long-form linear fiction from open-ended themes remains a major challenge for large language models, which frequently fail to guarantee g

BITS Pilani at SemEval-2026 Task 9: Structured Supervised Fine-Tuning with DPO Refinement for Polarization Detection

Model ReleasesDGX agent

arXiv:2604.11121v1 Announce Type: new Abstract: The POLAR SemEval-2026 Shared Task aims to detect online polarization and focuses on the classification and identification of multilingual, multicultura

BlasBench: An Open Benchmark for Irish Speech Recognition

Model ReleasesDGX agent

arXiv:2604.10736v1 Announce Type: new Abstract: No open Irish-specific benchmark compares end-user ASR systems under a shared Irish-aware evaluation protocol. To solve this, we release BlasBench, an o

BLUEmed: Retrieval-Augmented Multi-Agent Debate for Clinical Error Detection

Model ReleasesDGX agent

arXiv:2604.10389v1 Announce Type: new Abstract: Terminology substitution errors in clinical notes, where one medical term is replaced by a linguistically valid but clinically different term, pose a pe

BMdataset: A Musicologically Curated LilyPond Dataset

ResearchDGX agent

arXiv:2604.10628v1 Announce Type: cross Abstract: Symbolic music research has relied almost exclusively on MIDI-based datasets; text-based engraving formats such as LilyPond remain unexplored for musi

Both Ends Count! Just How Good are LLM Agents at 'Text-to-Big SQL'?

Model ReleasesDGX agent

arXiv:2602.21480v4 Announce Type: replace-cross Abstract: Text-to-SQL and Big Data are both extensively benchmarked fields, yet there is limited research that evaluates them jointly. In the real world

Bridging What the Model Thinks and How It Speaks: Self-Aware Speech Language Models for Expressive Speech Generation

Model ReleasesDGX agent

arXiv:2604.11424v1 Announce Type: new Abstract: Speech Language Models (SLMs) exhibit strong semantic understanding, yet their generated speech often sounds flat and fails to convey expressive intent,

CARINOX: Inference-time Scaling with Category-Aware Reward-based Initial Noise Optimization and Exploration

Model ReleasesDGX agent

arXiv:2509.17458v3 Announce Type: replace-cross Abstract: Text-to-image diffusion models, such as Stable Diffusion, can produce high-quality and diverse images but often fail to achieve compositional

CArtBench: Evaluating Vision-Language Models on Chinese Art Understanding, Interpretation, and Authenticity

Model ReleasesDGX agent

arXiv:2604.11632v1 Announce Type: new Abstract: We introduce CARTBENCH, a museum-grounded benchmark for evaluating vision-language models (VLMs) on Chinese artworks beyond short-form recognition and Q

Catalog-Native LLM: Speaking Item-ID Dialect with Less Entanglement for Recommendation

ResearchDGX agent

arXiv:2510.05125v2 Announce Type: replace Abstract: While collaborative filtering delivers predictive accuracy and efficiency, and Large Language Models (LLMs) enable expressive and generalizable reas

Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

Model ReleasesDGX agent

arXiv:2503.21380v3 Announce Type: replace Abstract: The rapid advancement of large reasoning models has saturated existing math benchmarks, underscoring the urgent need for more challenging evaluation

ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents

AgentsDGX agent

arXiv:2509.22830v3 Announce Type: replace Abstract: The growing deployment of large language model (LLM) based agents that interact with external environments has created new attack surfaces for adver

ChemPro: A Progressive Chemistry Benchmark for Large Language Models

Model ReleasesDGX agent

arXiv:2602.03108v3 Announce Type: replace Abstract: We introduce ChemPro, a progressive benchmark with 4100 natural language question-answer pairs in Chemistry, across 4 coherent sections of difficult

Claim2Vec: Embedding Fact-Check Claims for Multilingual Similarity and Clustering

SafetyDGX agent

arXiv:2604.09812v1 Announce Type: new Abstract: Recurrent claims present a major challenge for automated fact-checking systems designed to combat misinformation, especially in multilingual settings. W

ClaimDB: A Fact Verification Benchmark over Large Structured Data

Model ReleasesDGX agent

arXiv:2601.14698v2 Announce Type: replace Abstract: Real-world fact-checking often involves verifying claims grounded in structured data at scale. Despite substantial progress in fact-verification ben

CLSGen: A Dual-Head Fine-Tuning Framework for Joint Probabilistic Classification and Verbalized Explanation

Model ReleasesDGX agent

arXiv:2604.11801v1 Announce Type: new Abstract: With the recent progress of Large Language Models (LLMs), there is a growing interest in applying these models to solve complex and challenging problems

CodeComp: Structural KV Cache Compression for Agentic Coding

Local AiDGX agent

arXiv:2604.10235v1 Announce Type: new Abstract: Agentic code tasks such as fault localization and patch generation require processing long codebases under tight memory constraints, where the Key-Value

Comparative Analysis of Large Language Models in Healthcare

Model ReleasesDGX agent

arXiv:2604.10316v1 Announce Type: new Abstract: Background: Large Language Models (LLMs) are transforming artificial intelligence applications in healthcare due to their ability to understand, generat

CounterBench: Evaluating and Improving Counterfactual Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2502.11008v2 Announce Type: replace Abstract: Counterfactual reasoning is widely recognized as one of the most challenging and intricate aspects of causality in artificial intelligence. In this

Decomposing and Reducing Hidden Measurement Error in LLM Evaluation Pipelines

Model ReleasesDGX agent

arXiv:2604.11581v1 Announce Type: new Abstract: LLM evaluations drive which models get deployed, which safety standards get adopted, and which research conclusions get published. Yet these scores carr

DeCoVec: Building Decoding Space based Task Vector for Large Language Models via In-Context Learning

ResearchDGX agent

arXiv:2604.11129v1 Announce Type: new Abstract: Task vectors, representing directions in model or activation spaces that encode task-specific behaviors, have emerged as a promising tool for steering l

← Previous
1…119120121122123…128
Next →