AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlog
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Hardware

UCCI: Calibrated Uncertainty for Cost-Optimal LLM Cascade Routing

DGX agent

arXiv:2605.18796v1 Announce Type: cross Abstract: LLM cascades and model routing promise lower inference cost by sending easy queries to a small model and escalating hard ones to a large model, but mo

hardwarearxiv-cs-cl
20 May 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Unified Deployment-Aware Evaluation of Open Reasoning Language Models

DGX agent

arXiv:2604.07035v2 Announce Type: replace Abstract: Open reasoning language models are often compared under mixed sample sizes, partially standardized prompts, and accuracy-centered summaries, which m

model-releasesarxiv-cs-cl
20 May 2026
Research

What Are LLMs Doing to Scientific Communication? Measuring Changes in Writing Practices and Reading Experience

DGX agent

arXiv:2605.19936v1 Announce Type: new Abstract: Has the style of scientific communication changed due to the growing use of large language models in the writing process? We address this question in th

researcharxiv-cs-cl
20 May 2026
Research

Where Does Authorship Signal Emerge in Encoder-Based Language Models?

DGX agent

arXiv:2605.19908v1 Announce Type: new Abstract: Authorship attribution models fine-tuned with the same pretrained encoder, data, and loss can differ four-fold in performance depending only on their sc

researcharxiv-cs-cl
20 May 2026
Model Releases

XNote: Benchmarking Automated Community Notes Generation for Image-based Contextual Deception

DGX agent

arXiv:2603.22453v2 Announce Type: replace Abstract: Community Notes have emerged as an effective crowd-sourced mechanism for combating online deception on social media platforms. However, its reliance

model-releasesarxiv-cs-cl
20 May 2026
Model Releases

ZeroSearch: Incentivize the Search Capability of LLMs without Searching

DGX agent

arXiv:2505.04588v3 Announce Type: replace Abstract: Effective information searching is essential for enhancing the reasoning and generation capabilities of large language models (LLMs). Recent researc

model-releasesarxiv-cs-cl
20 May 2026
Model Releases

A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMDelta Integration into Upcycled MoE

DGX agent

arXiv:2605.18083v1 Announce Type: new Abstract: Expanding Large Language Models~(LLMs) to new languages is a costly endeavor, demanding extensive Continued Pre-Training~(CPT) and data-intensive alignm

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

A Pilot Benchmark for NL-to-FOL Translation in Planetary Exploration

DGX agent

arXiv:2605.17911v1 Announce Type: new Abstract: Future planetary exploration envisions autonomous robotic agents operating under severe communication constraints, without global positioning, and with

model-releasesarxiv-cs-cl
19 May 2026
Research

ACIL: Auto Chain of Thoughts for In-Context Learning

DGX agent

arXiv:2605.17088v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have shown that Chain-of-Thought (CoT) reasoning can substantially improve performance on complex reason

researcharxiv-cs-cl
19 May 2026
Safety

AI Agents May Always Fall for Prompt Injections

DGX agent

arXiv:2605.17634v1 Announce Type: cross Abstract: Prompt injection is the most critical vulnerability in deployed AI agents. Despite recent progress, we show that the prevailing defense paradigm (data

safetyarxiv-cs-cl
19 May 2026
Safety

AI Alignment Breaks at the Edge

DGX agent

arXiv:2602.20042v2 Announce Type: replace Abstract: General Alignment has improved average-case helpfulness and safety, but current alignment practice still rewards confident, single-turn responses. T

safetyarxiv-cs-cl
19 May 2026
Safety

AMATA: Adaptive Multi-Agent Trajectory Alignment for Knowledge-Intensive Question Answering

DGX agent

arXiv:2605.17352v1 Announce Type: new Abstract: Despite substantial advances in large language models (LLMs), generating factually consistent responses for knowledge-intensive question answering remai

safetyarxiv-cs-cl
19 May 2026
Research

Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades

DGX agent

arXiv:2605.17443v1 Announce Type: new Abstract: We analyze how automatic speech recognition (ASR) errors propagate through ASR-LLM cascades in Korean spoken question answering (SQA), focusing on downs

researcharxiv-cs-cl
19 May 2026
Model Releases

Ancient Greek to Modern Greek Machine Translation: A Novel Benchmark and Fine-Tuning Experiments on LLMs and NMT Models

DGX agent

arXiv:2605.18504v1 Announce Type: new Abstract: Machine Translation (MT) for Ancient Greek (AG) to Modern Greek (MG) is a low-resource task, constrained by the lack of large-scale, high-quality parall

model-releasesarxiv-cs-cl
19 May 2026
Safety

Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making

DGX agent

arXiv:2605.17228v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in high-stakes domains such as clinical decision support and medical documentation. However, the

safetyarxiv-cs-cl
19 May 2026
Research

AutoVecCoder: Teaching LLMs to Generate Explicitly Vectorized Code

DGX agent

arXiv:2605.17978v1 Announce Type: new Abstract: Vectorization via Single Instruction, Multiple Data (SIMD) architectures is a cornerstone of high-performance computing. To fully exploit hardware poten

researcharxiv-cs-cl
19 May 2026
Research

BELIEF: Structured Evidence Modeling and Uncertainty-Aware Fusion for Biomedical Question Answering

DGX agent

arXiv:2605.17435v1 Announce Type: new Abstract: Biomedical question answering often requires decisions from retrieved literature whose relevance, quality, and support for candidate answers are uneven.

researcharxiv-cs-cl
19 May 2026
Model Releases

Beyond Neural Incompatibility: Cross-Scale Knowledge Transfer in Language Models through Latent Semantic Alignment

DGX agent

arXiv:2510.24208v2 Announce Type: replace Abstract: Language Models (LMs) encode substantial knowledge in their parameters, yet it remains unclear how to transfer such knowledge in a fine-grained mann

model-releasesarxiv-cs-cl
19 May 2026
Applications

Beyond Sentiment Classification: A Generative Framework for Emotion Intensity Evaluation in Text

DGX agent

arXiv:2605.16613v1 Announce Type: new Abstract: We introduce a novel approach to emotion modeling that shifts the focus from identification to evaluation, addressing the limitations of discrete classi

applicationsarxiv-cs-cl
19 May 2026
Safety

Beyond the Final Actor: Modeling the Dual Roles of Creator and Editor for Fine-Grained LLM-Generated Text Detection

DGX agent

arXiv:2604.04932v3 Announce Type: replace Abstract: The misuse of large language models (LLMs) requires precise detection of synthetic text. Existing works mainly follow binary or ternary classificati

safetyarxiv-cs-cl
19 May 2026
Safety

Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech

DGX agent

arXiv:2605.17652v1 Announce Type: new Abstract: There are not enough established benchmarks for the task fo speech summarization. Creating new benchmarks demands human annotation, as LLMs could embed

safetyarxiv-cs-cl
19 May 2026
Research

Bridging the Gap: Converting Read Text to Conversational Dialogue

DGX agent

arXiv:2605.18001v1 Announce Type: new Abstract: In recent advancements within speech processing, converting read speech to conversational speech has gained significant attention. The primary challenge

researcharxiv-cs-cl
19 May 2026
Model Releases

Can LLMs Generate and Solve Linguistic Olympiad Puzzles?

DGX agent

arXiv:2509.21820v2 Announce Type: replace Abstract: In this paper, we introduce a combination of novel and exciting tasks: the solution and generation of linguistic puzzles. We focus on puzzles used i

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Closing the Gap at CRAC 2026: Two-Stage Adaptation for LLM-Based Multilingual Coreference Resolution

DGX agent

arXiv:2605.16984v1 Announce Type: new Abstract: We present our submission to the LLM track of the 2026 Computational Models of Reference, Anaphora and Coreference (CRAC 2026) shared task. With an aver

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection

DGX agent

arXiv:2605.16839v1 Announce Type: new Abstract: Chunked prefill has become a widely adopted serving strategy for long-context large language models, but efficient attention computation in this regime

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Compounding Disadvantage: Auditing Intersectional Bias in LLM-Generated Explanations Across Indian and American STEM Education

DGX agent

arXiv:2601.14506v3 Announce Type: replace-cross Abstract: Large language models are increasingly deployed in STEM education for personalized instruction and feedback across institutions in high- and l

model-releasesarxiv-cs-cl
19 May 2026
Safety

Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression

DGX agent

arXiv:2605.17304v1 Announce Type: cross Abstract: LLM context is not just tokens; it is a set of commitments. Long-running conversations accumulate goals, constraints, decisions, preferences, tool res

safetyarxiv-cs-cl
19 May 2026
Research

Confidence Geometry Reveals Trace-Level Correctness in Large Language Model Reasoning

DGX agent

arXiv:2605.16824v1 Announce Type: cross Abstract: Large language models (LLMs) generate not only reasoning text, but also token-level confidence trajectories that record how uncertainty evolves during

researcharxiv-cs-cl
19 May 2026
Local Ai

Constrained Code Generation with Discrete Diffusion

DGX agent

arXiv:2605.16829v1 Announce Type: new Abstract: Discrete diffusion models are a powerful, emerging paradigm for code generation. They construct programs through iterative refinement of partially corru

local-aiarxiv-cs-cl
19 May 2026
Safety

DISA: Offline Importance Sampling for Distribution-Matching LLM-RL

DGX agent

arXiv:2605.17295v1 Announce Type: cross Abstract: Modern reasoning agents are increasingly evaluated on their ability to generate multiple valid solution paths, plans, or tool-use traces for a given i

safetyarxiv-cs-cl
19 May 2026
Model Releases

Disentangling Ambiguity from Instability in Large Language Models: A Clinical Text-to-SQL Case Study

DGX agent

arXiv:2602.12015v2 Announce Type: replace Abstract: Deploying large language models for clinical Text-to-SQL requires distinguishing two qualitatively different causes of output diversity: (i) input a

model-releasesarxiv-cs-cl
19 May 2026
Safety

Do LLM Agents Mirror Socio-Cognitive Effects in Power-Asymmetric Conversations?

DGX agent

arXiv:2605.17694v1 Announce Type: new Abstract: Power differences shape human communication through well documented socio cognitive effects, including language coordination, pronoun usage, authority b

safetyarxiv-cs-cl
19 May 2026
Safety

Dual-Space Knowledge Distillation with Key-Query Matching for Large Language Models with Vocabulary Mismatch

DGX agent

arXiv:2603.22056v2 Announce Type: replace Abstract: Large language models (LLMs) achieve state-of-the-art (SOTA) performance across language tasks, but are costly to deploy due to their size and resou

safetyarxiv-cs-cl
19 May 2026
Research

E-PMQ: Expert-Guided Post-Merge Quantization with Merged-Weight Anchoring

DGX agent

arXiv:2605.16882v1 Announce Type: new Abstract: Low-resource deployment constraints have made model quantization essential for deploying neural networks while preserving performance. Meanwhile, model

researcharxiv-cs-cl
19 May 2026
Research

Early Stopping Chain-of-thoughts in Large Language Models

DGX agent

arXiv:2509.14004v2 Announce Type: replace Abstract: Reasoning large language models (LLMs) have demonstrated superior capacities in solving complicated problems by generating long chain-of-thoughts (C

researcharxiv-cs-cl
19 May 2026
Model Releases

Easier to Judge than to Find: Predicting In-Context Learning Success for Demonstration Selection

DGX agent

arXiv:2605.18512v1 Announce Type: new Abstract: In-context learning (ICL) is highly sensitive to which demonstrations appear in the prompt, but selecting them is expensive because the space of possibl

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Embodied Task Planning via Graph-Informed Action Generation with Large Language Models

DGX agent

arXiv:2601.21841v3 Announce Type: replace Abstract: While Large Language Models (LLMs) have demonstrated strong zero-shot reasoning capabilities, their deployment as embodied agents still faces fundam

model-releasesarxiv-cs-cl
19 May 2026
Research

Embracing Anisotropy: Turning Massive Activations into Interpretable Control Knobs for Large Language Models

DGX agent

arXiv:2603.00029v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit highly anisotropic internal representations, often characterized by massive activations, a phenomenon where a s

researcharxiv-cs-cl
19 May 2026
Agents

EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL

DGX agent

arXiv:2605.18703v1 Announce Type: new Abstract: Equipping LLMs with tool-use capabilities via Agentic Reinforcement Learning (Agentic RL) is bottlenecked by two challenges: the lack of scalable, robus

agentsarxiv-cs-cl
19 May 2026
Research

Evaluation Drift in LLM Personality Induction: Are We Moving the Goalpost?

DGX agent

arXiv:2605.16996v1 Announce Type: new Abstract: Can large language models reliably express a human-like personality, or are they merely mimicking surface cues without a stable underlying profile? To i

researcharxiv-cs-cl
19 May 2026
Model Releases

Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs

DGX agent

arXiv:2511.12710v2 Announce Type: replace Abstract: Automated red teaming frameworks for Large Language Models (LLMs) have become increasingly sophisticated, yet many still formulate attack optimizati

model-releasesarxiv-cs-cl
19 May 2026
Safety

Factual Inconsistencies in Multilingual Wikipedia Tables

DGX agent

arXiv:2507.18406v2 Announce Type: replace Abstract: Wikipedia serves as a globally accessible knowledge source with content in over 300 languages. Despite covering the same topics, the different versi

safetyarxiv-cs-cl
19 May 2026
Local Ai

FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing

DGX agent

arXiv:2605.17447v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have shown strong promise on Optical Character Recognition (OCR), yet the sheer number of visual tokens required to enco

local-aiarxiv-cs-cl
19 May 2026
Model Releases

FIM-LoRA: Task-Informative Rank Allocation for LoRA via Calibration-Time Gradient-Variance Estimation

DGX agent

arXiv:2605.16800v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) assigns a uniform rank to every adapted weight matrix - a practical convenience that ignores a fundamental reality: differe

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs

DGX agent

arXiv:2510.08886v3 Announce Type: replace Abstract: Going beyond simple text processing, financial auditing requires detecting semantic, structural, and numerical inconsistencies across large-scale di

model-releasesarxiv-cs-cl
19 May 2026
Research

Finding Sense in Nonsense with Generated Contexts: Perspectives from Humans and Language Models

DGX agent

arXiv:2602.11699v3 Announce Type: replace Abstract: Nonsensical and anomalous sentences have been instrumental in the development of computational models of semantic interpretation. A core challenge i

researcharxiv-cs-cl
19 May 2026
Model Releases

Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs

DGX agent

arXiv:2605.17558v1 Announce Type: cross Abstract: Training tool-calling agents requires large-scale trajectory data with verifiable labels, yet existing approaches either synthesize environments that

model-releasesarxiv-cs-cl
19 May 2026
Local Ai

FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers

DGX agent

arXiv:2605.17231v1 Announce Type: cross Abstract: Activation steering methods modify intermediate representations of language models to control output behavior, but universally assume the activation s

local-aiarxiv-cs-cl
19 May 2026
← Previous
1…8889909192…162
Next →