AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
21 Apr 2026

CBR-to-SQL: Rethinking Retrieval-based Text-to-SQL using Case-based Reasoning in the Healthcare Domain

ApplicationsDGX agent

arXiv:2603.05569v2 Announce Type: replace-cross Abstract: Extracting insights from Electronic Health Record (EHR) databases often requires SQL expertise, creating a barrier for clinical decision-makin

CBRS: Cognitive Blood Request System with Bilingual Dataset and Dual-Layer Filtering for Multi-Platform Social Streams

Model ReleasesDGX agent

arXiv:2604.16665v1 Announce Type: new Abstract: Urgent blood donation seeking posts and messages on social media often go unnoticed due to the overwhelming volume of daily communications. Traditional

CFMS: Towards Explainable and Fine-Grained Chinese Multimodal Sarcasm Detection Benchmark

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2604.16372v1 Announce Type: new Abstract: Multimodal sarcasm detection has recently garnered significant attention. However, existing benchmarks suffer from coarse-grained annotations and limite

Characterizing Model-Native Skills

SafetyDGX agent

arXiv:2604.17614v1 Announce Type: cross Abstract: Skills are a natural unit for describing what a language model can do and how its behavior can be changed. However, existing characterizations rely on

CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and Ideation

ResearchDGX agent

arXiv:2505.20779v5 Announce Type: replace Abstract: A hallmark of human innovation is recombination -- the creation of novel ideas by integrating elements from existing concepts and mechanisms. In thi

CLAG: Adaptive Memory Organization via Agent-Driven Clustering for Small Language Model Agents

AgentsDGX agent

arXiv:2603.15421v2 Announce Type: replace Abstract: Large language model agents heavily rely on external memory to support knowledge reuse and complex reasoning tasks. Yet most memory systems store ex

ClawEnvKit: Automatic Environment Generation for Claw-Like Agents

Model ReleasesDGX agent

arXiv:2604.18543v1 Announce Type: cross Abstract: Constructing environments for training and evaluating claw-like agents remains a manual, human-intensive process that does not scale. We argue that wh

Clinical Note Bloat Reduction for Efficient LLM Use

ResearchDGX agent

arXiv:2604.16364v1 Announce Type: cross Abstract: Health systems are rapidly deploying large language models (LLMs) that use clinical notes for clinical decision support applications. However, modern

Closing the Modality Reasoning Gap for Speech Large Language Models

SafetyDGX agent

arXiv:2601.05543v2 Announce Type: replace Abstract: Although Speech Large Language Models have achieved notable progress, a substantial modality reasoning gap remains: their reasoning performance on s

CoAct: Co-Active LLM Preference Learning with Human-AI Synergy

ResearchDGX agent

arXiv:2604.17501v1 Announce Type: new Abstract: Learning from preference-based feedback has become an effective approach for aligning LLMs across diverse tasks. However, high-quality human-annotated p

CodePivot: Bootstrapping Multilingual Transpilation in LLMs via Reinforcement Learning without Parallel Corpora

Model ReleasesDGX agent

arXiv:2604.18027v1 Announce Type: cross Abstract: Transpilation, or code translation, aims to convert source code from one programming language (PL) to another. It is beneficial for many downstream ap

CoDial: Interpretable Task-Oriented Dialogue Systems Through Dialogue Flow Alignment

Model ReleasesDGX agent

arXiv:2506.02264v3 Announce Type: replace Abstract: Building Task-Oriented Dialogue (TOD) systems that generalize across different tasks remains a challenging problem. Data-driven approaches often str

Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations

Model ReleasesDGX agent

arXiv:2507.20409v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting helps models think step by step. But naive CoT breaks down in visually grounded social tasks, where models must per

Cognitive Policy-Driven LLM for Diagnosis and Intervention of Cognitive Distortions in Emotional Support Conversation

SafetyDGX agent

arXiv:2604.17178v1 Announce Type: new Abstract: Emotional Support Conversation (ESC) plays a critical role in mental health assistance by providing accessible psychological support in real-world appli

Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

ApplicationsDGX agent

arXiv:2506.01732v2 Announce Type: replace Abstract: Large Language Models (LLMs) are pre-trained on large data from different sources and domains. These datasets often contain trillions of tokens, inc

Comparing Human and Large Language Model Interpretation of Implicit Information

ResearchDGX agent

arXiv:2604.17085v1 Announce Type: new Abstract: The interpretation of implicit meanings is an integral aspect of human communication. However, this framework may not transfer to interactions with Larg

ComPASS: Towards Personalized Agentic Social Support via Tool-Augmented Companionship

Model ReleasesDGX agent

arXiv:2604.18356v1 Announce Type: new Abstract: Developing compassionate interactive systems requires agents to not only understand user emotions but also provide diverse, substantive support. While r

Compositional Steering of Large Language Models with Steering Tokens

ApplicationsDGX agent

arXiv:2601.05062v2 Announce Type: replace Abstract: Deploying LLMs in real-world applications requires controllable output that satisfies multiple desiderata at the same time. While existing work exte

Concurrent Criterion Validation of a Validity Screen for LLM Confidence Signals via Selective Prediction

Model ReleasesDGX agent

arXiv:2604.17716v1 Announce Type: new Abstract: The validity screen (Cacioli, 2026d, 2026e) classifies LLM confidence signals as Valid, Indeterminate, or Invalid. We test whether these classifications

Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning

HardwareDGX agent

arXiv:2412.00069v3 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) has garnered significant attention for its ability to scale up neural networks while utilizing the same or even fewer

Contrastive Analysis of Linguistic Representations in Large Language Model Outputs through Structured Synthetic Data Generation and Abstracted N-gram Associations

SafetyDGX agent

arXiv:2604.17398v1 Announce Type: new Abstract: We present a methodological framework to discover linguistic and discursive patterns associated to different social groups through contrastive synthetic

Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks

ResearchDGX agent

arXiv:2604.17761v1 Announce Type: cross Abstract: Interpretability tools are increasingly used to analyze failures of Large Language Models (LLMs), yet prior work largely focuses on short prompts or t

ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling

TutorialsDGX agent

arXiv:2510.08878v3 Announce Type: replace-cross Abstract: Text-to-audio (TTA) generation with fine-grained control signals, e.g., precise timing control or intelligible speech content, has been explor

Copy-as-Decode: Grammar-Constrained Parallel Prefill for LLM Editing

HardwareDGX agent

arXiv:2604.18170v1 Announce Type: new Abstract: LLMs edit text and code by autoregressively regenerating the full output, even when most tokens appear verbatim in the input. We study Copy-as-Decode, a

Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining

Model ReleasesDGX agent

arXiv:2604.17633v1 Announce Type: new Abstract: Large language models exhibit impressive cross-lingual capabilities. However, prior work analyzes this phenomenon through isolated factors and at sparse

COSEARCH: Joint Training of Reasoning and Document Ranking via Reinforcement Learning for Agentic Search

SafetyDGX agent

arXiv:2604.17555v1 Announce Type: cross Abstract: Agentic search -- the task of training agents that iteratively reason, issue queries, and synthesize retrieved information to answer complex questions

Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR

ResearchDGX agent

arXiv:2603.07084v2 Announce Type: replace-cross Abstract: Reward hacking is a form of misalignment in which models overoptimize proxy rewards without genuinely solving the underlying task. Precisely m

Creating ConLangs to Probe the Metalinguistic Grammatical Knowledge of LLMs

Model ReleasesDGX agent

arXiv:2510.07591v3 Announce Type: replace Abstract: We present a system that uses LLMs as a tool in the development of Constructed Languages -- ConLangs, which we call IASC (Interactive Agentic System

CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit

Model ReleasesDGX agent

arXiv:2510.06133v2 Announce Type: replace Abstract: Diffusion large language models (dLLMs) generate text through iterative denoising. In commonly adopted parallel decoding schemes, each step confirms

CRISP: Compressing Redundancy in Chain-of-Thought via Intrinsic Saliency Pruning

SafetyDGX agent

arXiv:2604.17297v1 Announce Type: new Abstract: Long Chain-of-Thought (CoT) reasoning is pivotal for the success of recent reasoning models but suffers from high computational overhead and latency. Wh

CROC: Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks

Model ReleasesDGX agent

arXiv:2505.11314v2 Announce Type: replace-cross Abstract: The assessment of evaluation metrics (meta-evaluation) is crucial for determining the suitability of existing metrics in text-to-image (T2I) g

Cross-Family Speculative Decoding for Polish Language Models on Apple~Silicon: An Empirical Evaluation of Bielik~11B with UAG-Extended MLX-LM

Model ReleasesDGX agent

arXiv:2604.16368v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by using a small draft model to propose k candidate tokens for a target model to verify. While effective

Crowded in B-Space: Calibrating Shared Directions for LoRA Merging

ApplicationsDGX agent

arXiv:2604.16826v1 Announce Type: new Abstract: Merging separately trained LoRA adapters is a practical alternative to joint multi-task training, but it often hurts performance. Existing methods usual

CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction

ApplicationsDGX agent

arXiv:2604.16742v1 Announce Type: cross Abstract: Scientists have long sought to accurately predict outcomes of real-world events before they happen. Can AI systems do so more reliably? We study this

Culinary Crossroads: A RAG Framework for Enhancing Diversity in Cross-Cultural Recipe Adaptation

ResearchDGX agent

arXiv:2507.21934v2 Announce Type: replace Abstract: In cross-cultural recipe adaptation, the goal is not only to ensure cultural appropriateness and retain the original dish's essence, but also to pro

Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts

SafetyDGX agent

arXiv:2604.18091v1 Announce Type: new Abstract: Recent multimodal large language models have shown promising ability in generating humorous captions for images, yet they still lack stable control over

DART: Mitigating Harm Drift in Difference-Aware LLMs via Distill-Audit-Repair Training

Model ReleasesDGX agent

arXiv:2604.16845v1 Announce Type: new Abstract: Large language models (LLMs) tuned for safety often avoid acknowledging demographic differences, even when such acknowledgment is factually correct (e.g

Data Compressibility Quantifies LLM Memorization

ResearchDGX agent

arXiv:2507.06056v4 Announce Type: replace Abstract: Large Language Models (LLMs) are known to memorize portions of their training data, sometimes even reproduce content verbatim when prompted appropri

Data Mixing for Large Language Models Pretraining: A Survey and Outlook

ResearchDGX agent

arXiv:2604.16380v1 Announce Type: new Abstract: Large language models (LLMs) rely on pretraining on massive and heterogeneous corpora, where training data composition has a decisive impact on training

Decisive: Guiding User Decisions with Optimal Preference Elicitation from Unstructured Documents

ResearchDGX agent

arXiv:2604.18122v1 Announce Type: new Abstract: Decision-making is a cognitively intensive task that requires synthesizing relevant information from multiple unstructured sources, weighing competing f

Decoupling the Effect of Chain-of-Thought Reasoning: A Human Label Variation Perspective

SafetyDGX agent

arXiv:2601.03154v2 Announce Type: replace Abstract: Reasoning-tuned LLMs utilizing long Chain-of-Thought (CoT) excel at single-answer tasks, yet their ability to model Human Label Variation--which req

Defragmenting Language Models: An Interpretability-based Approach for Vocabulary Expansion

ResearchDGX agent

arXiv:2604.16656v1 Announce Type: new Abstract: All languages are equal; when it comes to tokenization, some are more equal than others. Tokens are the hidden currency that dictate the cost and latenc

DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models

ResearchDGX agent

arXiv:2604.17709v1 Announce Type: new Abstract: Existing works on large language model (LLM) decomposition mainly focus on improving performance on downstream tasks, but they ignore the poor parallel

Demystifying the unreasonable effectiveness of online alignment methods

SafetyDGX agent

arXiv:2604.17207v1 Announce Type: cross Abstract: Iterative alignment methods based on purely greedy updates are remarkably effective in practice, yet existing theoretical guarantees of (O(log T)) KL-

Depth Registers Unlock W4A4 on SwiGLU: A Reader/Generator Decomposition

Model ReleasesDGX agent

arXiv:2604.18128v1 Announce Type: new Abstract: We study post-training W4A4 quantization in a controlled 300M-parameter SwiGLU decoder-only language model trained on 5B tokens of FineWeb-Edu, and ask

Designing Explainable Conversational Agentic Systems for Guarani Speakers

AgentsDGX agent

arXiv:2603.05743v3 Announce Type: replace Abstract: Although artificial intelligence (AI) and Human-Computer Interaction (HCI) systems are often presented as universal solutions, their design remains

Detecting Alarming Student Verbal Responses using Text and Audio Classifier

SafetyDGX agent

arXiv:2604.16717v1 Announce Type: new Abstract: This paper addresses a critical safety gap in the use Automated Verbal Response Scoring (AVRS). We present a novel hybrid framework for troubled student

Detecting LLM-Generated Spam Reviews by Integrating Language Model Embeddings and Graph Neural Network

Model ReleasesDGX agent

arXiv:2510.01801v2 Announce Type: replace Abstract: The rise of large language models (LLMs) has enabled the generation of highly persuasive spam reviews that closely mimic human writing. These review

Diagnosing LLM-based Rerankers in Cold-Start Recommender Systems: Coverage, Exposure and Practical Mitigations

SafetyDGX agent

arXiv:2604.16318v1 Announce Type: cross Abstract: Large language models (LLMs) and cross-encoder rerankers have gained attention for improving recommender systems, particularly in cold-start scenarios

DiffCoT: Diffusion-styled Chain-of-Thought Reasoning in LLMs

SafetyDGX agent

arXiv:2601.03559v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) reasoning improves multi-step mathematical problem solving in large language models but remains vulnerable to exposure bias a

Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks

Local AiDGX agent

arXiv:2604.18510v1 Announce Type: cross Abstract: Open-weight language models can be rendered unsafe through several distinct interventions, but the resulting models may differ substantially in capabi

Diversity Collapse in Multi-Agent LLM Systems: Structural Coupling and Collective Failure in Open-Ended Idea Generation

AgentsDGX agent

arXiv:2604.18005v1 Announce Type: cross Abstract: Multi-agent systems (MAS) are increasingly used for open-ended idea generation, driven by the expectation that collective interaction will broaden the

Do LLMs Encode Functional Importance of Reasoning Tokens?

ResearchDGX agent

arXiv:2601.03066v2 Announce Type: replace Abstract: Large language models solve complex tasks by generating long reasoning chains, achieving higher accuracy at the cost of increased computational cost

Do LLMs Use Cultural Knowledge Without Being Told? A Multilingual Evaluation of Implicit Pragmatic Adaptation

SafetyDGX agent

arXiv:2604.17718v1 Announce Type: new Abstract: Many benchmarks show that large language models can answer direct questions about culture. We study a different question: do they also change how they s

DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-Completion

Model ReleasesDGX agent

arXiv:2604.18257v1 Announce Type: cross Abstract: Query auto-completion (QAC) has been widely studied in the context of web search, yet remains underexplored for in-document search, which we term DocQ

Document-as-Image Representations Fall Short for Scientific Retrieval

Model ReleasesDGX agent

arXiv:2604.18508v1 Announce Type: cross Abstract: Many recent document embedding models are trained on document-as-image representations, embedding rendered pages as images rather than the underlying

Does Welsh media need a review? Detecting bias in Nation.Cymru's political reporting

SafetyDGX agent

arXiv:2604.17628v1 Announce Type: new Abstract: Wales' political landscape has been marked by growing accusations of bias in Welsh media. This paper takes the first computational step toward testing t

Domain-oriented RAG Assessment (DoRA): Synthetic Benchmarking for RAG-based Question Answering on Defense Documents

Model ReleasesDGX agent

arXiv:2604.17943v1 Announce Type: new Abstract: Open-domain RAG benchmarks over public corpora can overestimate deployment performance due to pretraining overlap and weak attribution requirements. We

Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models

AgentsDGX agent

arXiv:2510.07248v3 Announce Type: replace Abstract: Small language models (SLMs) enable scalable tool-augmented multi-agent systems where multiple SLMs handle subtasks orchestrated by a powerful coord

DORA Explorer: Improving the Exploration Ability of LLMs Without Training

Model ReleasesDGX agent

arXiv:2604.17244v1 Announce Type: new Abstract: Despite the rapid progress, LLMs for sequential decision-making (i.e., LLM agents) still struggle to produce diverse outputs. This leads to insufficient

← Previous
1…105106107108109…129
Next →