AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
28 Apr 2026

A BERTology View of LLM Orchestrations: Token- and Layer-Selective Probes for Efficient Single-Pass Classification

Model ReleasesDGX agent

arXiv:2601.13288v2 Announce Type: replace Abstract: Production LLM systems often rely on separate models for safety and other classification-heavy steps, increasing latency, VRAM footprint, and operat

A Comparative analysis of Layer-wise Representational Capacity in AR and Diffusion LLMs

SafetyDGX agent

arXiv:2603.07475v2 Announce Type: replace Abstract: Autoregressive (AR) language models build representations incrementally via left-to-right prediction, while diffusion language models (dLLMs) are tr

A Large-Scale, Cross-Disciplinary Corpus of Systematic Reviews

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.22864v1 Announce Type: cross Abstract: Existing benchmarks for systematic reviewing remain limited either in scale or in disciplinary coverage, with some collections comprising only a modes

A Multi-Dimensional Audit of Politically Aligned Large Language Models

SafetyDGX agent

arXiv:2604.24429v1 Announce Type: new Abstract: As the application of Large Language Models (LLMs) spreads across various industries, there are increasing concerns about the potential for their misuse

A Survey on Split Learning for LLM Fine-Tuning: Models, Systems, and Privacy Optimizations

ResearchDGX agent

arXiv:2604.24468v1 Announce Type: cross Abstract: Fine-tuning unlocks large language models (LLMs) for specialized applications, but its high computational cost often puts it out of reach for resource

AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards

Model ReleasesDGX agent

arXiv:2604.22840v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong potential in agentic tasks, particularly in slide generation. However, slide generation poses a

AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking

AgentsDGX agent

arXiv:2604.23581v1 Announce Type: cross Abstract: Agentic systems that chain reasoning, tool use, and synthesis into multi-step workflows are entering production, yet prevailing evaluation practices l

AI use in American newspapers is widespread, uneven, and rarely disclosed

Local AiDGX agent

arXiv:2510.18774v4 Announce Type: replace Abstract: AI is rapidly transforming journalism, but the extent of its use in published newspaper articles remains unclear. We address this gap by auditing a

ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis

SafetyDGX agent

arXiv:2404.10141v2 Announce Type: replace-cross Abstract: Text-to-image (T2I) models have achieved remarkable progress in high-quality image synthesis, yet most benchmarks rely on simple, self-contain

AP-BMM: Approximating Capability-Efficiency Pareto Sets of LLMs via Asynchronous Prior-guided Bayesian Model Merging

ResearchDGX agent

arXiv:2512.09972v5 Announce Type: replace-cross Abstract: Navigating the capability--efficiency trade-off in Large Language Models (LLMs) requires approximating a high-quality Pareto set. Existing mod

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs

Model ReleasesDGX agent

arXiv:2604.22937v1 Announce Type: new Abstract: Verification is becoming central to both reinforcement-learning-based training and inference-time control of large language models (LLMs). Yet current v

Benchmarking Testing in Automated Theorem Proving

Model ReleasesDGX agent

arXiv:2604.23698v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have shown promise in formal theorem proving, yet evaluating semantic correctness remains challenging. E

Beyond Local vs. External: A Game-Theoretic Framework for Trustworthy Knowledge Acquisition

Model ReleasesDGX agent

arXiv:2604.23413v1 Announce Type: new Abstract: Cloud-hosted Large Language Models (LLMs) offer unmatched reasoning capabilities and dynamic knowledge, yet submitting raw queries to these external ser

BiMol-Diff: A Unified Diffusion Framework for Molecular Generation and Captioning

ResearchDGX agent

arXiv:2604.24089v1 Announce Type: new Abstract: Bridging molecular structures and natural language is essential for controllable design. Autoregressive models struggle with long-range dependencies, wh

Bridging Reasoning and Action: Hybrid LLM-RL Framework for Efficient Cross-Domain Task-Oriented Dialogue

SafetyDGX agent

arXiv:2604.23345v1 Announce Type: new Abstract: Cross-domain task-oriented dialogue requires reasoning over implicit and explicit feasibility constraints while planning long-horizon, multi-turn action

Bridging the Domain Divide: Supervised vs. Zero-Shot Clinical Section Segmentation from MIMIC-III to Obstetrics

ApplicationsDGX agent

arXiv:2602.17513v2 Announce Type: replace Abstract: Clinical free-text notes contain vital patient information. They are structured into labelled sections; recognizing these sections has been shown to

BRIEF-Pro: Universal Context Compression with Short-to-Long Synthesis for Fast and Accurate Multi-Hop Reasoning

ResearchDGX agent

arXiv:2510.13799v2 Announce Type: replace Abstract: As retrieval-augmented generation (RAG) tackles complex tasks, increasingly expanded contexts offer richer information, but at the cost of higher la

Can Compact Language Models Search Like Agents? Distillation-Guided Policy Optimization for Preserving Agentic RAG Capabilities

SafetyDGX agent

arXiv:2508.20324v4 Announce Type: replace Abstract: Reinforcement Learning has emerged as a dominant post-training approach to elicit agentic RAG behaviors such as search and planning from language mo

Can Humans Detect AI? Mining Textual Signals of AI-Assisted Writing Under Varying Scrutiny Conditions

ResearchDGX agent

arXiv:2604.23471v1 Announce Type: cross Abstract: This study asks whether the threat of AI detection changes how people write with AI, and whether other people can tell the difference. In a two-phase

Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination

Model ReleasesDGX agent

arXiv:2604.24690v1 Announce Type: new Abstract: While Large Language Models (LLMs) have increasingly assisted in historical tasks such as text processing, their capacity for professional-level histori

Can You Make It Sound Like You? Post-Editing LLM-Generated Text for Personal Style

ResearchDGX agent

arXiv:2604.24444v1 Announce Type: new Abstract: Despite the growing use of large language models (LLMs) for writing tasks, users may hesitate to rely on LLMs when personal style is important. Post-edi

ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question Answering

ResearchDGX agent

arXiv:2510.13312v2 Announce Type: replace Abstract: We present ChatR1, a reasoning framework based on reinforcement learning (RL) for conversational question answering (CQA). Reasoning plays an import

Chinese-SkillSpan: A Span-Level Dataset for ESCO-Aligned Competency Extraction from Chinese Job Ads

Model ReleasesDGX agent

arXiv:2604.23009v1 Announce Type: new Abstract: Job Skill Named Entity Recognition (JobSkillNER) aims to automatically extract key skill information from large-scale job posting data, which is importa

CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era

Model ReleasesDGX agent

arXiv:2602.23452v2 Announce Type: replace Abstract: Scientific research relies on accurate citation for attribution and integrity, yet large language models (LLMs) introduce a new risk: fabricated ref

ComplianceNLP: Knowledge-Graph-Augmented RAG for Multi-Framework Regulatory Gap Detection

Model ReleasesDGX agent

arXiv:2604.23585v1 Announce Type: new Abstract: Financial institutions must track over 60,000 regulatory events annually, overwhelming manual compliance teams; the industry has paid over USD 300 billi

Contextual Linear Activation Steering of Language Models

ResearchDGX agent

arXiv:2604.24693v1 Announce Type: new Abstract: Linear activation steering is a powerful approach for eliciting the capabilities of large language models and specializing their behavior using limited

ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents

AgentsDGX agent

arXiv:2604.23069v1 Announce Type: new Abstract: Large language model (LLM) agents often struggle in long-context interactions. As the agent accumulates more interaction history, context management app

CRISP: Persistent Concept Unlearning via Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2508.13650v3 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in real-world applications, the need to selectively remove unwanted knowledge while preser

Culture-Aware Machine Translation in Large Language Models: Benchmarking and Investigation

ResearchDGX agent

arXiv:2604.24361v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance in general machine translation, yet their ability in culture-aware scenarios remains poorl

DARC-CLIP: Dynamic Adaptive Refinement with Cross-Attention for Meme Understanding

Model ReleasesDGX agent

arXiv:2604.23214v1 Announce Type: new Abstract: Memes convey meaning through the interaction of visual and textual signals, often combining humor, irony, and offense in subtle ways. Detecting harmful

DeepTaxon: An Interpretable Retrieval-Augmented Multimodal Framework for Unified Species Identification and Discovery

Model ReleasesDGX agent

arXiv:2604.24029v1 Announce Type: cross Abstract: Identifying species in biology among tens of thousands of visually similar taxa while discovering unknown species in open-world environments remains a

Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models

ResearchDGX agent

arXiv:2601.02455v2 Announce Type: replace-cross Abstract: Deploying Automatic Speech Recognition (ASR) models on memory-constrained edge devices requires aggressive low-bit weight quantization. Layer-

Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer

Model ReleasesDGX agent

arXiv:2604.24302v1 Announce Type: new Abstract: Mechanistic interpretability has made it possible to localize circuits underlying specific behaviors in language models, but existing methods are expens

Domain Fine-Tuning vs. Retrieval-Augmented Generation for Medical Multiple-Choice Question Answering: A Controlled Comparison at the 4B-Parameter Scale

Model ReleasesDGX agent

arXiv:2604.23801v1 Announce Type: new Abstract: Practitioners deploying small open-weight large language models (LLMs) for medical question answering face a recurring design choice: invest in a domain

DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based Agents

SafetyDGX agent

arXiv:2604.24320v1 Announce Type: new Abstract: Large language model (LLM) agents that follow the sequential 'reason-then-act' paradigm have achieved superior performance in many complex tasks.However

DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute

AgentsDGX agent

arXiv:2604.23815v1 Announce Type: new Abstract: Scientific Deep Research (DR) agents answer user queries by synthesizing research papers into multi-section reports. User feedback can improve their uti

DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models

ResearchDGX agent

arXiv:2505.13975v4 Announce Type: replace Abstract: While Large Reasoning Models (LRMs) have demonstrated success in complex reasoning tasks through long chain-of-thought (CoT) reasoning, their infere

DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack

ApplicationsDGX agent

arXiv:2512.16182v2 Announce Type: replace-cross Abstract: With the rapid development of cloud-based services, large language models have become increasingly accessible through various web platforms. H

Efficient Rationale-based Retrieval: On-policy Distillation from Generative Rerankers based on JEPA

SafetyDGX agent

arXiv:2604.23336v1 Announce Type: cross Abstract: Unlike traditional fact-based retrieval, rationale-based retrieval typically necessitates cross-encoding of query-document pairs using large language

EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving

Model ReleasesDGX agent

arXiv:2604.22851v1 Announce Type: cross Abstract: While Vision-Language Models (VLMs) have advanced highlevel reasoning in autonomous driving, their ability to ground this reasoning in the underlying

Evaluating Large Language Models on Computer Science University Exams in Data Structures

Model ReleasesDGX agent

arXiv:2604.23347v1 Announce Type: new Abstract: We present a comprehensive evaluation of Large Language Models (LLMs) on Computer Science (CS) Data Structure examination questions. Our work introduces

Evaluating Temporal Consistency in Multi-Turn Language Models

Model ReleasesDGX agent

arXiv:2604.23051v1 Announce Type: new Abstract: Language models are increasingly deployed in interactive settings where users reason about facts over time rather than in isolation. In such scenarios,

Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models

ResearchDGX agent

arXiv:2510.02629v3 Announce Type: replace Abstract: Context utilisation, the ability of Language Models (LMs) to incorporate relevant information from the provided context when generating responses, r

Evaluation of Pose Estimation Systems for Sign Language Translation

ResearchDGX agent

arXiv:2604.24609v1 Announce Type: new Abstract: Many sign language translation (SLT) systems operate on pose sequences instead of raw video to reduce input dimensionality, improve portability, and par

Evolve: A Persistent Knowledge Lifecycle for Small Language Models

Model ReleasesDGX agent

arXiv:2604.23424v1 Announce Type: cross Abstract: Evolve pairs a small local language model with a persistent, teacher-compiled knowledge store -- refined through sleep consolidation and usage-driven

EXCEEDS: Extracting Complex Events via Nugget-based Grid Modeling in Scientific Domain

Model ReleasesDGX agent

arXiv:2406.14075v2 Announce Type: replace Abstract: It is crucial to understand a specific domain by events. Extensive event extraction research has been conducted in many domains such as news, financ

Explaining Sources of Uncertainty in Automated Fact-Checking

ResearchDGX agent

arXiv:2505.17855v2 Announce Type: replace Abstract: Understanding sources of a model's uncertainty regarding its predictions is crucial for effective human-AI collaboration. Prior work proposes using

Factual and Edit-Sensitive Graph-to-Sequence Generation via Graph-Aware Adaptive Noising

Local AiDGX agent

arXiv:2604.24104v1 Announce Type: new Abstract: Fine-tuned autoregressive models for graph-to-sequence generation (G2S) often struggle with factual grounding and edit sensitivity. To tackle these issu

Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning Perspective

TutorialsDGX agent

arXiv:2604.23267v1 Announce Type: new Abstract: Large language models (LLMs) operate in two fundamental learning modes - fine-tuning (FT) and in-context learning (ICL) - raising key questions about wh

Food4All: A Multi-Agent Framework for Real-time Free Food Discovery with Integrated Nutritional Metadata

AgentsDGX agent

arXiv:2510.18289v2 Announce Type: replace Abstract: Food insecurity remains a persistent public health emergency in the United States, tightly interwoven with chronic disease, mental illness, and opio

For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMs

Model ReleasesDGX agent

arXiv:2508.10180v3 Announce Type: replace Abstract: Data valuation is essential for enhancing the transparency and accountability of large language models (LLMs) and vision-language models (VLMs). How

Gated Tree Cross-Attention for Checkpoint-Compatible Syntax Injection in Decoder-Only LLMs

ResearchDGX agent

arXiv:2602.15846v2 Announce Type: replace Abstract: Decoder-only large language models achieve strong broad performance but are brittle to minor grammatical perturbations, undermining reliability for

Generating Place-Based Compromises Between Two Points of View

Model ReleasesDGX agent

arXiv:2604.24536v1 Announce Type: new Abstract: Large Language Models (LLMs) excel academically but struggle with social intelligence tasks, such as creating good compromises. In this paper, we presen

GraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMs

Model ReleasesDGX agent

arXiv:2604.23626v1 Announce Type: new Abstract: LLM routing has achieved promising results in integrating the strengths of diverse models while balancing efficiency and performance. However, to suppor

HalalBench: A Multilingual OCR Benchmark for Food Packaging Ingredient Extraction

Model ReleasesDGX agent

arXiv:2604.22754v1 Announce Type: cross Abstract: No standardized benchmark exists for evaluating OCR on food packaging, despite its critical role in automated halal food verification. Existing benchm

HeadRouter: Dynamic Head-Weight Routing for Task-Adaptive Audio Token Pruning in Large Audio Language Models

ResearchDGX agent

arXiv:2604.23717v1 Announce Type: cross Abstract: Recent large audio language models (LALMs) demonstrate remarkable capabilities in processing extended multi-modal sequences, yet incur high inference

Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance

Local AiDGX agent

arXiv:2604.23318v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) performs coarse-grained credit assignment in reinforcement learning with verifiable rewards (RLVR) by assignin

How Sensitive Are Safety Benchmarks to Judge Configuration Choices?

Model ReleasesDGX agent

arXiv:2604.24074v1 Announce Type: new Abstract: Safety benchmarks such as HarmBench rely on LLM judges to classify model responses as harmful or safe, yet the judge configuration, namely the combinati

Implicit Framing in Obstetric Counseling Notes: A Grounded LLM Pipeline on a VBAC-Eligible Cohort

ApplicationsDGX agent

arXiv:2604.23059v1 Announce Type: new Abstract: Clinical framing -- the linguistic manner in which clinical information is presented -- can influence patient understanding and decision-making, with im

In-depth Analysis of Graph-based RAG in a Unified Framework

ResearchDGX agent

arXiv:2503.04338v2 Announce Type: replace-cross Abstract: Graph-based Retrieval-Augmented Generation (RAG) has proven effective in integrating external knowledge into large language models (LLMs), imp

← Previous
1…96979899100…129
Next →