AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
28 May 2026

Adaptive Cost-Efficient Evaluation for Reliable Patent Claim Generation

Model ReleasesDGX agent

arXiv:2604.04295v3 Announce Type: replace Abstract: Automated patent claim validation demands low error tolerance. However, existing approaches face a rigidity-resource dilemma: lightweight encoders c

Addressing Pitfalls in Auditing Practices of Automatic Speech Recognition Technologies: A Case Study of People with Aphasia

ApplicationsDGX agent

arXiv:2506.08846v3 Announce Type: replace-cross Abstract: Automatic Speech Recognition (ASR) systems' growing use warrants robust auditing approaches to ensure equitable transcription quality, especia

AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2512.17375v2 Announce Type: replace-cross Abstract: LLM-as-a-Judge systems supply the reward signal in modern RLHF and RLVR pipelines, but their binary verdict reduces to a single linear readout

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning

SafetyDGX agent

arXiv:2605.28774v1 Announce Type: new Abstract: Vision-language models with extended reasoning succeed on complex problems, but many real-world problems require external tools that internal reasoning

Agentic Separation Logic Specification Synthesis

Model ReleasesDGX agent

arXiv:2605.27531v1 Announce Type: cross Abstract: Specification synthesis, the task of automatically inferring formal specifications from program implementations and natural language, is important for

Agents that Matter: Optimizing Multi-Agent LLMs via Removal-Based Attribution

SafetyDGX agent

arXiv:2605.27621v1 Announce Type: cross Abstract: As multi-agent systems (MAS) become increasingly complex, identifying the contributions of individual agents is critical for system optimization. Howe

AI Research Agents Narrow Scientific Exploration

AgentsDGX agent

arXiv:2605.27905v1 Announce Type: new Abstract: AI research agents can now generate research ideas, design experiments, run code, and draft papers, raising the possibility of large-scale AI-assisted s

An Evolutionary Approach for Designing Stable and Highly Expressible Low-Immunogenicity Therapeutic mRNA Sequences

Local AiDGX agent

arXiv:2605.27986v1 Announce Type: new Abstract: Messenger RNA (mRNA) sequences as therapeutics require optimized design to ensure efficient translation, structural stability, and minimal immunogenicit

Analyzing Cancer Patients' Experiences with Embedding-based Topic Modeling and LLMs

ApplicationsDGX agent

arXiv:2601.12154v2 Announce Type: replace Abstract: This study investigates the use of neural topic modeling and LLMs to uncover meaningful themes from patient storytelling data, to offer insights tha

Analyzing Quality-Latency-Resource Trade-offs in a Technical Documentation RAG Assistant Using LoRA Adaptation

Model ReleasesDGX agent

arXiv:2605.28222v1 Announce Type: new Abstract: We study quality-latency-resource trade-offs in a documentation-grounded retrieval-augmented generation (RAG) system that uses Low-Rank Adaptation (LoRA

Are We Truly Innovating? A Qualitative and Quantitative Study of Originality in AI Research Papers

ResearchDGX agent

arXiv:2602.06054v3 Announce Type: replace Abstract: Assessing originality in AI research is arguably the most consequential yet least reliable step in peer review. Reviewer judgments of originality re

Argument Quality Assessment with Large Language Models: A Pairwise Bradley-Terry Approach

Model ReleasesDGX agent

arXiv:2605.28313v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in tasks related to reasoning and judgment. However, assessing the quality of arg

Ask Now, Use Later: Benchmarking the Proactivity Gap in Long-Lived LLM Agents

Model ReleasesDGX agent

arXiv:2605.28108v1 Announce Type: new Abstract: A long-lived LLM agent, such as OpenClaw, earns its value by acting on a user's preferences and constraints across sessions, not just the current reques

Assessing Factual Music Comprehension in Large Audio Language Models

Model ReleasesDGX agent

arXiv:2511.05550v2 Announce Type: replace-cross Abstract: Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio

ATLAS: All-round Testing of Long-context Abilities across Scales

Model ReleasesDGX agent

arXiv:2605.28079v1 Announce Type: new Abstract: Long-context language models now advertise context windows up to millions of tokens, yet evaluations typically report a single length or a narrow task f

Attention Projection Mixing with Exogenous Anchors

ResearchDGX agent

arXiv:2601.08131v4 Announce Type: replace Abstract: Cross-layer reuse of early attention projections can improve optimization and data efficiency, but it creates a structural conflict: the first layer

Auditing Stance Asymmetry in Generative Explanations

SafetyDGX agent

arXiv:2605.27988v1 Announce Type: new Abstract: Bias evaluation for language models has made substantial progress on bounded comparisons, such as overt derogation, stereotype association, or label-sen

BEAR: Budgeted Evidence Allocation for Multi-Document Reasoning

ApplicationsDGX agent

arXiv:2601.18116v2 Announce Type: replace Abstract: We argue that multi-document reasoning is constrained not only by how much text a model can read, but also by how limited query-time evidence budget

Benchmarking and Mechanistic Analysis of Vision-Language Models for Cross-Depiction Assembly Instruction Alignment

Model ReleasesDGX agent

arXiv:2604.00913v2 Announce Type: replace-cross Abstract: 2D assembly diagrams are often abstract and hard to follow, creating a need for intelligent assistants that can monitor progress, detect error

Better heads do not guarantee better binarized constituency parsing

ResearchDGX agent

arXiv:2605.28131v1 Announce Type: new Abstract: We revisit punctuation-aware tree binarization for constituency parsing and ask whether dependency-induced headedness improves binary parser supervision

Beyond Chunk-Local Extraction: Cross-Chunk Graph Augmentation for GraphRAG

ResearchDGX agent

arXiv:2605.28004v1 Announce Type: new Abstract: GraphRAG extends retrieval-augmented generation by organizing corpora as explicit knowledge graphs, enabling graph-based retrieval for complex question

Beyond Input Understanding: Diagnosing Multilingual Mathematical Reasoning with Directed Acyclic Trace Graphs

ResearchDGX agent

arXiv:2605.27715v1 Announce Type: new Abstract: Large reasoning models (LRMs) achieve strong mathematical reasoning performance in English, but remain much less reliable in many low- and medium-resour

Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents

Model ReleasesDGX agent

arXiv:2605.28465v1 Announce Type: new Abstract: Divergent thinking is a core dimension of creativity, yet existing evaluations of Large Language Models (LLMs) treat them as single-turn text generation

Beyond pass@k: Redundancy-Aware RLVR for Multi-Sample Code Generation

ResearchDGX agent

arXiv:2605.28022v1 Announce Type: new Abstract: LLMs for code generation are commonly evaluated in repeated-sampling settings using Pass@k, where multiple candidate programs are executed against unit

Boundary Suppression Asymmetry in Post-trained Assistants: Over-expansion as a Controllability Cost

SafetyDGX agent

arXiv:2605.27969v1 Announce Type: new Abstract: Post-trained language-model assistants are often optimized to avoid under-answering, encouraging complete, helpful, cautious, and proactive responses. W

Breaking the Script Barrier: Enabling Automatic Alignment for PoS-based ASR Error Analysis in Non-Latin Scripts

SafetyDGX agent

arXiv:2605.28438v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) systems are commonly evaluated using aggregate metrics such as Word Error Rate (WER), which do not capture the lingui

Building Community-Centred NLP Resources for Puno Quechua

Model ReleasesDGX agent

arXiv:2605.28253v1 Announce Type: new Abstract: The preservation of under-resourced languages requires digital tools and resources shaped by and for their speakers. We present the first dedicated ASR

CALM-IT: Generating Realistic Long-Form Motivational Interviewing Dialogues with Dual-Actor Conversational Dynamics Tracking

ResearchDGX agent

arXiv:2601.10085v2 Announce Type: replace Abstract: Therapeutic dialogue is not a sequence of isolated responses: client goals, motivation, resistance, and therapeutic alliance evolve over time. Yet c

Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

Model ReleasesDGX agent

arXiv:2510.05291v2 Announce Type: replace Abstract: As Large Language Models (LLMs) develop stronger multilingual capabilities, their sensitivity to culturally diverse entities becomes increasingly im

Can Hallucinations Be Useful? Solving Multi-Hop Questions With SLMs By Chaining System-I/II Reasoning

ResearchDGX agent

arXiv:2605.27596v1 Announce Type: new Abstract: Recently, there has been increased interest in Small Language Models (SLMs), which are fast, show good performance, and have lower hardware demands than

Can Large Language Models Handle Discourse Particles? A Case Study of Colloquial Malay

Model ReleasesDGX agent

arXiv:2605.28782v1 Announce Type: new Abstract: Discourse particles, such as extit{well} and extit{kind of}, are crucial components that enable LLMs to ``speak'' more like humans. They are used to con

Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence?

ResearchDGX agent

arXiv:2605.28778v1 Announce Type: new Abstract: LLMs' linguistically expressed confidence should faithfully reflect their intrinsic uncertainty. While recent work shows LLMs struggle to use epistemic

CAREF: Calibration-Aware Regularization for Explanation Faithfulness Without Rationale Supervision

Model ReleasesDGX agent

arXiv:2605.27835v1 Announce Type: cross Abstract: We introduce CAREF, a parameter-efficient fine-tuning framework that jointly optimizes predictive accuracy and explanation faithfulness via calibratio

Chain-based Adaptive Reconfiguration Over Lattices for Hallucination Reduction

AgentsDGX agent

arXiv:2605.27706v1 Announce Type: new Abstract: We introduce CAROL (Chain-based Adaptive Reconfiguration Over Lattices), a probabilistic framework for test-time hallucination reduction in large langua

Challenges in Explaining Pretrained Clinical Text Classifiers

ResearchDGX agent

arXiv:2605.28060v1 Announce Type: new Abstract: Explaining the predictions of neural models in clinical NLP remains a significant challenge, especially for complex tasks involving long, unstructured m

Chinese Word Boundary Recovery through Character Alignment Projection

Model ReleasesDGX agent

arXiv:2605.28128v1 Announce Type: new Abstract: Chinese word segmentation is especially fragile in non-standard text, where language learner errors and other character-level divergences disrupt the wo

CIRF: Tokenizing Chain-of-Thoughts into Reusable Functional Units for Efficient Latent Reasoning in Large Language Models

SafetyDGX agent

arXiv:2605.28292v1 Announce Type: new Abstract: Implicit Chain-of-Thought (CoT) reduces the inference cost of large language models by internalizing the explicit rationales. However, existing approach

ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs

Model ReleasesDGX agent

arXiv:2603.02097v5 Announce Type: replace Abstract: Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in l

ClinicalAgents: Multi-Agent Orchestration for Clinical Decision Making with Dual-Memory

Model ReleasesDGX agent

arXiv:2603.26182v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have demonstrated potential in healthcare, they often struggle with the complex, non-linear reasoning required fo

ClinicalEncoder26AM: A Multlilingual Diagnosable ColBERT Model; Evidences from the MultiClinNER Shared Task

Local AiDGX agent

arXiv:2605.28521v1 Announce Type: new Abstract: ClinicalEncoder26AM is a multilingual Diagnosable ColBERT for clinical and biomedical texts, which aligns at multiple levels its token-level semantic wi

Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests

Model ReleasesDGX agent

arXiv:2605.28734v1 Announce Type: cross Abstract: A general-purpose language model that answers a harmful question returns text; a coding model that complies with a malicious request can return a work

CodeGENCAT: Generative Computerized Adaptive Testing for Open-ended Coding Problems

SafetyDGX agent

arXiv:2602.20020v2 Announce Type: replace Abstract: Existing Computerized Adaptive Testing (CAT) frameworks typically select questions based on the predicted likelihood that the student will answer co

Comonadic Morphophonology: A Compositional Framework for Context-Dependent Morphological Rules in Finnish

ResearchDGX agent

arXiv:2605.28484v1 Announce Type: new Abstract: Composing finite-state transducers (FSTs) for context-dependent morphophonological rules -- consonant gradation, vowel harmony, possessive suffix assimi

ConRAG: Consensus-Driven Multi-View Retrieval for Multi-Hop Question Answering

Model ReleasesDGX agent

arXiv:2605.28093v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) has emerged as a promising paradigm for enhancing large language models (LLMs) on multi-hop question answering (QA)

ConvMemory: A Lightweight Learned Memory Reranker, a Negative Attribution Result, and a Research-Preview Conflict Editor

Model ReleasesDGX agent

arXiv:2605.28062v1 Announce Type: new Abstract: We describe ConvMemory, a small 3.6M-parameter learned reranker for conversational long-term memory retrieval, trained with cross-encoder teacher superv

Decoupling Skeleton and Flesh: Efficient Multimodal Table Reasoning with Disentangled Alignment and Structure-aware Guidance

Local AiDGX agent

arXiv:2602.03491v2 Announce Type: replace-cross Abstract: Reasoning over table images remains challenging for Large Vision-Language Models (LVLMs) due to complex layouts and tightly coupled structure-

DisasterBench: Benchmarking LLM Planning under Typed Tool Interface Constraints

Model ReleasesDGX agent

arXiv:2605.27957v1 Announce Type: new Abstract: Disasters cause severe societal impacts, demanding rapid coordination of heterogeneous AI tools, from satellite analysis to flood prediction and damage

Disentangling Language Roles in Multilingual LLM Task Execution

Model ReleasesDGX agent

arXiv:2605.27649v1 Announce Type: new Abstract: Multilingual LLMs are increasingly used when instruction, source content, and required response languages do not coincide. Existing benchmarks have expa

DRTriton: Large-Scale Synthetic Data Driven Reinforcement Learning for Triton Kernel Generation

Model ReleasesDGX agent

arXiv:2603.21465v2 Announce Type: replace Abstract: Developing efficient CUDA kernels is a fundamental yet challenging task in the generative AI industry. Recent research leverages Large Language Mode

Escape the Language Prior: Mitigating Late-Stage Modality Collapse in Audio Reasoning via Modality-Aware Policy Optimization

SafetyDGX agent

arXiv:2605.27741v1 Announce Type: new Abstract: Audio and omni-modal large language models exhibit impressive cross-modal reasoning capabilities. However, applying standard reinforcement learning post

Evaluating the Generation Capabilities of Large Chinese Language Models

ResearchDGX agent

arXiv:2308.04823v5 Announce Type: replace Abstract: This paper unveils CG-Eval, the first-ever comprehensive and automated evaluation framework designed for assessing the generative capabilities of la

Explanation Generation for Contradiction Reconciliation with LLMs

ResearchDGX agent

arXiv:2603.22735v2 Announce Type: replace Abstract: Existing NLP work commonly treats contradictions as errors to be resolved by choosing which statements to accept or discard. Yet a key aspect of hum

FABSVer: Faster Training and Better Self-Verification for LLM Mathematical Reasoning

SafetyDGX agent

arXiv:2605.28389v1 Announce Type: new Abstract: While large language models have made significant progress in mathematical reasoning, they remain unreliable at judging the correctness of their own sol

FEA-SLT: A Gloss-Free End-to-End Framework for Facial-Expression-Aware Sign Language Translation

ResearchDGX agent

arXiv:2601.03549v2 Announce Type: replace-cross Abstract: Sign Language Translation (SLT) is a challenging cross-modal task requiring joint modeling of manual articulations and non-manual signals. Exi

FinBoardBench: Benchmarking Dynamic Wealth Management and Strategic Financial Reasoning of LLMs via Board Game Simulations

ApplicationsDGX agent

arXiv:2605.27896v1 Announce Type: new Abstract: Recently, large language models (LLMs) have achieved superior performance in static financial reasoning and simple dynamic trading tasks. However, exist

Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language Models

ResearchDGX agent

arXiv:2510.17620v2 Announce Type: replace Abstract: Large language models may encode sensitive information or outdated knowledge that needs to be removed, to ensure responsible and compliant model res

Formula-One Prompting: A Composable Equation-First Prefix for Applied Mathematics

TutorialsDGX agent

arXiv:2601.19302v3 Announce Type: replace Abstract: This paper introduces Formula Prompting (FP) and Formula-One Prompting (F-1), two single-call methods that elicit governing equations before solving

Framing Matters: Addressing Framing Sensitivity in Decision-Making through Behaviorally-Grounded Value Alignment

Model ReleasesDGX agent

arXiv:2605.28188v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in high-stakes decision-making settings such as legal reasoning, where consistency under factuall

GeneralThinker: Domain-General Reasoning through Likelihood-Guided Answer-Conditioned Optimization

SafetyDGX agent

arXiv:2605.27934v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards improves language model reasoning, but its reliance on domain-specific verifiers, sparse outcome rewards,

GRADE: Generalizable Reasoning-Aware Dialogue Evaluation for AI Tutors

ResearchDGX agent

arXiv:2605.27866v1 Announce Type: new Abstract: Evaluating AI tutor responses requires more than factual correctness: tutors must identify mistakes, locate errors, provide guidance, and offer actionab

← Previous
1…5556575859…129
Next →