AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation

DGX agent

arXiv:2604.25235v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used as automated judges for multimodal systems, yet their scores provide no indication of reliability.

model-releasesarxiv-cs-cl
29 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution

DGX agent

arXiv:2604.03472v2 Announce Type: replace Abstract: Co-evolutionary self-play, where one language model generates problems and another solves them, promises autonomous curriculum learning without huma

safetyarxiv-cs-cl
29 Apr 2026
Safety

Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation

DGX agent

arXiv:2511.21517v2 Announce Type: replace Abstract: Unlike text, speech conveys information about the speaker, such as gender, through acoustic cues like pitch. This gives rise to modality-specific bi

safetyarxiv-cs-cl
29 Apr 2026
Research

VOYAGER: A Training Free Approach for Generating Diverse Datasets using LLMs

DGX agent

arXiv:2512.12072v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly being used to generate synthetic datasets for the evaluation and training of downstream models. Howeve

researcharxiv-cs-cl
29 Apr 2026
Model Releases

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

DGX agent

arXiv:2604.25591v1 Announce Type: cross Abstract: Recent audio-aware large language models (ALLMs) have demonstrated strong capabilities across diverse audio understanding and reasoning tasks, but the

model-releasesarxiv-cs-cl
29 Apr 2026
Research

What Makes Good Instruction-Tuning Data? An In-Context Learning Perspective

DGX agent

arXiv:2604.25132v1 Announce Type: new Abstract: Instruction-tuning datasets often contain substantial redundancy and low-quality samples, necessitating effective data selection methods. We propose an

researcharxiv-cs-cl
29 Apr 2026
Model Releases

When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs

DGX agent

arXiv:2510.07499v2 Announce Type: replace Abstract: Recent Long-Context Language Models (LCLMs) can process hundreds of thousands of tokens in a single prompt, enabling new opportunities for knowledge

model-releasesarxiv-cs-cl
29 Apr 2026
Hardware

WhisperPipe: A Resource-Efficient Streaming Architecture for Real-Time Automatic Speech Recognition

DGX agent

arXiv:2604.25611v1 Announce Type: new Abstract: Real-time automatic speech recognition (ASR) systems face a fundamental trade-off between transcription accuracy and computational efficiency, particula

hardwarearxiv-cs-cl
29 Apr 2026
Research

Why Do LLM-based Web Agents Fail? A Hierarchical Planning Perspective

DGX agent

arXiv:2603.14248v2 Announce Type: replace-cross Abstract: Large language model (LLM) web agents are increasingly used for web navigation but remain far from human reliability on realistic, long-horizo

researcharxiv-cs-cl
29 Apr 2026
Research

Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models

DGX agent

arXiv:2604.25011v1 Announce Type: new Abstract: Reinforcement learning (RL)-based post-training often improves the reasoning performance of large language models (LLMs) beyond the training domain, whi

researcharxiv-cs-cl
29 Apr 2026
Research

Wiki Dumps to Training Corpora: South Slavic Case

DGX agent

arXiv:2604.25384v1 Announce Type: new Abstract: This paper presents a methodology for transforming raw Wikimedia dumps into quality textual corpora for seven South Slavic languages. The work is divide

researcharxiv-cs-cl
29 Apr 2026
Model Releases

A Benchmark Suite of Reddit-Derived Datasets for Mental Health Detection

DGX agent

arXiv:2604.23458v1 Announce Type: new Abstract: The growing availability of online support groups has opened up new windows to study mental health through natural language processing (NLP). However, i

model-releasesarxiv-cs-cl
28 Apr 2026
Model Releases

A BERTology View of LLM Orchestrations: Token- and Layer-Selective Probes for Efficient Single-Pass Classification

DGX agent

arXiv:2601.13288v2 Announce Type: replace Abstract: Production LLM systems often rely on separate models for safety and other classification-heavy steps, increasing latency, VRAM footprint, and operat

model-releasesarxiv-cs-cl
28 Apr 2026
Safety

A Comparative analysis of Layer-wise Representational Capacity in AR and Diffusion LLMs

DGX agent

arXiv:2603.07475v2 Announce Type: replace Abstract: Autoregressive (AR) language models build representations incrementally via left-to-right prediction, while diffusion language models (dLLMs) are tr

safetyarxiv-cs-cl
28 Apr 2026
Research

A Large-Scale, Cross-Disciplinary Corpus of Systematic Reviews

DGX agent

arXiv:2604.22864v1 Announce Type: cross Abstract: Existing benchmarks for systematic reviewing remain limited either in scale or in disciplinary coverage, with some collections comprising only a modes

researcharxiv-cs-cl
28 Apr 2026
Safety

A Multi-Dimensional Audit of Politically Aligned Large Language Models

DGX agent

arXiv:2604.24429v1 Announce Type: new Abstract: As the application of Large Language Models (LLMs) spreads across various industries, there are increasing concerns about the potential for their misuse

safetyarxiv-cs-cl
28 Apr 2026
Research

A Survey on Split Learning for LLM Fine-Tuning: Models, Systems, and Privacy Optimizations

DGX agent

arXiv:2604.24468v1 Announce Type: cross Abstract: Fine-tuning unlocks large language models (LLMs) for specialized applications, but its high computational cost often puts it out of reach for resource

researcharxiv-cs-cl
28 Apr 2026
Model Releases

AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards

DGX agent

arXiv:2604.22840v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong potential in agentic tasks, particularly in slide generation. However, slide generation poses a

model-releasesarxiv-cs-cl
28 Apr 2026
Agents

AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking

DGX agent

arXiv:2604.23581v1 Announce Type: cross Abstract: Agentic systems that chain reasoning, tool use, and synthesis into multi-step workflows are entering production, yet prevailing evaluation practices l

agentsarxiv-cs-cl
28 Apr 2026
Local Ai

AI use in American newspapers is widespread, uneven, and rarely disclosed

DGX agent

arXiv:2510.18774v4 Announce Type: replace Abstract: AI is rapidly transforming journalism, but the extent of its use in published newspaper articles remains unclear. We address this gap by auditing a

local-aiarxiv-cs-cl
28 Apr 2026
Safety

ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis

DGX agent

arXiv:2404.10141v2 Announce Type: replace-cross Abstract: Text-to-image (T2I) models have achieved remarkable progress in high-quality image synthesis, yet most benchmarks rely on simple, self-contain

safetyarxiv-cs-cl
28 Apr 2026
Research

AP-BMM: Approximating Capability-Efficiency Pareto Sets of LLMs via Asynchronous Prior-guided Bayesian Model Merging

DGX agent

arXiv:2512.09972v5 Announce Type: replace-cross Abstract: Navigating the capability--efficiency trade-off in Large Language Models (LLMs) requires approximating a high-quality Pareto set. Existing mod

researcharxiv-cs-cl
28 Apr 2026
Model Releases

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs

DGX agent

arXiv:2604.22937v1 Announce Type: new Abstract: Verification is becoming central to both reinforcement-learning-based training and inference-time control of large language models (LLMs). Yet current v

model-releasesarxiv-cs-cl
28 Apr 2026
Model Releases

Benchmarking Testing in Automated Theorem Proving

DGX agent

arXiv:2604.23698v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have shown promise in formal theorem proving, yet evaluating semantic correctness remains challenging. E

model-releasesarxiv-cs-cl
28 Apr 2026
Model Releases

Beyond Local vs. External: A Game-Theoretic Framework for Trustworthy Knowledge Acquisition

DGX agent

arXiv:2604.23413v1 Announce Type: new Abstract: Cloud-hosted Large Language Models (LLMs) offer unmatched reasoning capabilities and dynamic knowledge, yet submitting raw queries to these external ser

model-releasesarxiv-cs-cl
28 Apr 2026
Research

BiMol-Diff: A Unified Diffusion Framework for Molecular Generation and Captioning

DGX agent

arXiv:2604.24089v1 Announce Type: new Abstract: Bridging molecular structures and natural language is essential for controllable design. Autoregressive models struggle with long-range dependencies, wh

researcharxiv-cs-cl
28 Apr 2026
Safety

Bridging Reasoning and Action: Hybrid LLM-RL Framework for Efficient Cross-Domain Task-Oriented Dialogue

DGX agent

arXiv:2604.23345v1 Announce Type: new Abstract: Cross-domain task-oriented dialogue requires reasoning over implicit and explicit feasibility constraints while planning long-horizon, multi-turn action

safetyarxiv-cs-cl
28 Apr 2026
Applications

Bridging the Domain Divide: Supervised vs. Zero-Shot Clinical Section Segmentation from MIMIC-III to Obstetrics

DGX agent

arXiv:2602.17513v2 Announce Type: replace Abstract: Clinical free-text notes contain vital patient information. They are structured into labelled sections; recognizing these sections has been shown to

applicationsarxiv-cs-cl
28 Apr 2026
Research

BRIEF-Pro: Universal Context Compression with Short-to-Long Synthesis for Fast and Accurate Multi-Hop Reasoning

DGX agent

arXiv:2510.13799v2 Announce Type: replace Abstract: As retrieval-augmented generation (RAG) tackles complex tasks, increasingly expanded contexts offer richer information, but at the cost of higher la

researcharxiv-cs-cl
28 Apr 2026
Safety

Can Compact Language Models Search Like Agents? Distillation-Guided Policy Optimization for Preserving Agentic RAG Capabilities

DGX agent

arXiv:2508.20324v4 Announce Type: replace Abstract: Reinforcement Learning has emerged as a dominant post-training approach to elicit agentic RAG behaviors such as search and planning from language mo

safetyarxiv-cs-cl
28 Apr 2026
Research

Can Humans Detect AI? Mining Textual Signals of AI-Assisted Writing Under Varying Scrutiny Conditions

DGX agent

arXiv:2604.23471v1 Announce Type: cross Abstract: This study asks whether the threat of AI detection changes how people write with AI, and whether other people can tell the difference. In a two-phase

researcharxiv-cs-cl
28 Apr 2026
Model Releases

Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination

DGX agent

arXiv:2604.24690v1 Announce Type: new Abstract: While Large Language Models (LLMs) have increasingly assisted in historical tasks such as text processing, their capacity for professional-level histori

model-releasesarxiv-cs-cl
28 Apr 2026
Research

Can You Make It Sound Like You? Post-Editing LLM-Generated Text for Personal Style

DGX agent

arXiv:2604.24444v1 Announce Type: new Abstract: Despite the growing use of large language models (LLMs) for writing tasks, users may hesitate to rely on LLMs when personal style is important. Post-edi

researcharxiv-cs-cl
28 Apr 2026
Research

ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question Answering

DGX agent

arXiv:2510.13312v2 Announce Type: replace Abstract: We present ChatR1, a reasoning framework based on reinforcement learning (RL) for conversational question answering (CQA). Reasoning plays an import

researcharxiv-cs-cl
28 Apr 2026
Model Releases

Chinese-SkillSpan: A Span-Level Dataset for ESCO-Aligned Competency Extraction from Chinese Job Ads

DGX agent

arXiv:2604.23009v1 Announce Type: new Abstract: Job Skill Named Entity Recognition (JobSkillNER) aims to automatically extract key skill information from large-scale job posting data, which is importa

model-releasesarxiv-cs-cl
28 Apr 2026
Model Releases

CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era

DGX agent

arXiv:2602.23452v2 Announce Type: replace Abstract: Scientific research relies on accurate citation for attribution and integrity, yet large language models (LLMs) introduce a new risk: fabricated ref

model-releasesarxiv-cs-cl
28 Apr 2026
Model Releases

ComplianceNLP: Knowledge-Graph-Augmented RAG for Multi-Framework Regulatory Gap Detection

DGX agent

arXiv:2604.23585v1 Announce Type: new Abstract: Financial institutions must track over 60,000 regulatory events annually, overwhelming manual compliance teams; the industry has paid over USD 300 billi

model-releasesarxiv-cs-cl
28 Apr 2026
Research

Contextual Linear Activation Steering of Language Models

DGX agent

arXiv:2604.24693v1 Announce Type: new Abstract: Linear activation steering is a powerful approach for eliciting the capabilities of large language models and specializing their behavior using limited

researcharxiv-cs-cl
28 Apr 2026
Agents

ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents

DGX agent

arXiv:2604.23069v1 Announce Type: new Abstract: Large language model (LLM) agents often struggle in long-context interactions. As the agent accumulates more interaction history, context management app

agentsarxiv-cs-cl
28 Apr 2026
Model Releases

CRISP: Persistent Concept Unlearning via Sparse Autoencoders

DGX agent

arXiv:2508.13650v3 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in real-world applications, the need to selectively remove unwanted knowledge while preser

model-releasesarxiv-cs-cl
28 Apr 2026
Research

Culture-Aware Machine Translation in Large Language Models: Benchmarking and Investigation

DGX agent

arXiv:2604.24361v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance in general machine translation, yet their ability in culture-aware scenarios remains poorl

researcharxiv-cs-cl
28 Apr 2026
Model Releases

DARC-CLIP: Dynamic Adaptive Refinement with Cross-Attention for Meme Understanding

DGX agent

arXiv:2604.23214v1 Announce Type: new Abstract: Memes convey meaning through the interaction of visual and textual signals, often combining humor, irony, and offense in subtle ways. Detecting harmful

model-releasesarxiv-cs-cl
28 Apr 2026
Model Releases

DeepTaxon: An Interpretable Retrieval-Augmented Multimodal Framework for Unified Species Identification and Discovery

DGX agent

arXiv:2604.24029v1 Announce Type: cross Abstract: Identifying species in biology among tens of thousands of visually similar taxa while discovering unknown species in open-world environments remains a

model-releasesarxiv-cs-cl
28 Apr 2026
Research

Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models

DGX agent

arXiv:2601.02455v2 Announce Type: replace-cross Abstract: Deploying Automatic Speech Recognition (ASR) models on memory-constrained edge devices requires aggressive low-bit weight quantization. Layer-

researcharxiv-cs-cl
28 Apr 2026
Model Releases

Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer

DGX agent

arXiv:2604.24302v1 Announce Type: new Abstract: Mechanistic interpretability has made it possible to localize circuits underlying specific behaviors in language models, but existing methods are expens

model-releasesarxiv-cs-cl
28 Apr 2026
Model Releases

Domain Fine-Tuning vs. Retrieval-Augmented Generation for Medical Multiple-Choice Question Answering: A Controlled Comparison at the 4B-Parameter Scale

DGX agent

arXiv:2604.23801v1 Announce Type: new Abstract: Practitioners deploying small open-weight large language models (LLMs) for medical question answering face a recurring design choice: invest in a domain

model-releasesarxiv-cs-cl
28 Apr 2026
Safety

DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based Agents

DGX agent

arXiv:2604.24320v1 Announce Type: new Abstract: Large language model (LLM) agents that follow the sequential 'reason-then-act' paradigm have achieved superior performance in many complex tasks.However

safetyarxiv-cs-cl
28 Apr 2026
Agents

DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute

DGX agent

arXiv:2604.23815v1 Announce Type: new Abstract: Scientific Deep Research (DR) agents answer user queries by synthesizing research papers into multi-section reports. User feedback can improve their uti

agentsarxiv-cs-cl
28 Apr 2026
← Previous
1…120121122123124…161
Next →