AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
15 May 2026

Cross-Linguistic Transcription and Phonological Representation in the Huitongguanxi Huayiyiyu

ResearchDGX agent

arXiv:2605.14480v1 Announce Type: new Abstract: Purpose: This study investigates the transcription principles underlying Huitongguanxi Huayiyiyu (HHY), a series of multilingual glossaries compiled by

Distribution Corrected Offline Data Distillation for Large Language Models

SafetyDGX agent

arXiv:2605.14071v1 Announce Type: new Abstract: Distilling reasoning traces from strong large language models into smaller ones is a promising route to improve intelligence in resource-constrained set

Do Composed Image Retrieval Benchmarks Require Multimodal Composition?

ResearchDGX agent

arXiv:2605.14787v1 Announce Type: cross Abstract: Composed Image Retrieval (CIR) is a multimodal retrieval task where a query consists of a reference image and a textual modification, and the goal is


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Do Reasoning LLMs Refuse What They Infer in Long Contexts?

SafetyDGX agent

arXiv:2602.08874v2 Announce Type: replace Abstract: Long-context LLMs can infer objectives that are not stated explicitly. This capability is useful for reasoning over documents, code, retrieved evide

Does Local News Stay Local?: Online Content Shifts in Sinclair-Acquired Stations

ResearchDGX agent

arXiv:2510.07060v2 Announce Type: replace Abstract: Local news stations are often considered to be reliable sources of non-politicized information, particularly local concerns that residents care abou

DT-Transformer: A Foundation Model for Disease Trajectory Prediction on a Real-world Health System

ApplicationsDGX agent

arXiv:2605.14227v1 Announce Type: cross Abstract: Accurate disease trajectory prediction is critical for early intervention, resource allocation, and improving long-term outcomes. While electronic hea

Dual Hierarchical Dialogue Policy Learning for Legal Inquisitive Conversational Agents

SafetyDGX agent

arXiv:2605.14057v1 Announce Type: new Abstract: Most existing dialogue systems are user-driven, primarily designed to fulfill user requests. However, in many critical real-world scenarios, a conversat

EndPrompt: Efficient Long-Context Extension via Terminal Anchoring

Model ReleasesDGX agent

arXiv:2605.14589v1 Announce Type: new Abstract: Extending the context window of large language models typically requires training on sequences at the target length, incurring quadratic memory and comp

Energy-Regularized Sequential Model Editing on Hyperspheres

ApplicationsDGX agent

arXiv:2510.01172v3 Announce Type: replace Abstract: Large language models (LLMs) require constant updates to remain aligned with evolving real-world knowledge. Model editing offers a lightweight alter

FactNet: A Billion-Scale Knowledge Graph for Multilingual Factual Grounding

ResearchDGX agent

arXiv:2602.03417v2 Announce Type: replace Abstract: Large language models hallucinate factual claims and struggle to ground their outputs in retrievable evidence, particularly in non-English languages

Factorization-Error-Free Discrete Diffusion Language Model via Speculative Decoding

ResearchDGX agent

arXiv:2605.14305v1 Announce Type: new Abstract: Discrete diffusion language models improve generation efficiency through parallel token prediction, but standard X_0 prediction methods introduce factor

Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution

Model ReleasesDGX agent

arXiv:2605.15138v1 Announce Type: cross Abstract: Standard unlearning evaluations measure behavioral suppression in full precision, immediately after training, despite every deployed language model be

From Scenes to Elements: Multi-Granularity Evidence Retrieval for Verifiable Multimodal RAG

Model ReleasesDGX agent

arXiv:2605.15019v1 Announce Type: new Abstract: Multimodal Retrieval-Augmented Generation (RAG) systems retrieve evidence at coarse granularities (entire images or scenes), creating a mismatch with fi

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

Model ReleasesDGX agent

arXiv:2605.15104v1 Announce Type: new Abstract: Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified

Geometry-Aware Decoding with Wasserstein-Regularized Truncation and Mass Penalties for Large Language Models

ResearchDGX agent

arXiv:2602.10346v2 Announce Type: replace Abstract: Large language models (LLMs) must balance diversity and creativity against logical coherence in open-ended generation. Existing truncation-based sam

GradShield: Alignment Preserving Finetuning

SafetyDGX agent

arXiv:2605.14194v1 Announce Type: new Abstract: Large Language Models (LLMs) pose a significant risk of safety misalignment after finetuning, as models can be compromised by both explicitly and implic

GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations

Model ReleasesDGX agent

arXiv:2605.14498v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly serve as personal assistants and workplace collaborators, where their utility depends on memory systems t

Ideology Prediction of German Political Texts

SafetyDGX agent

arXiv:2605.14352v1 Announce Type: new Abstract: Elections represent a crucial milestone in a nation's ongoing development. To better understand the political rhetoric from various movements, ranging f

Is Grep All You Need? How Agent Harnesses Reshape Agentic Search

Model ReleasesDGX agent

arXiv:2605.15184v1 Announce Type: new Abstract: Recent advances in Large Language Model (LLM) agents have enabled complex agentic workflows where models autonomously retrieve information, call tools,

It Takes Two: Your GRPO Is Secretly DPO

ResearchDGX agent

arXiv:2510.00977v3 Announce Type: replace-cross Abstract: GRPO has emerged as a prominent reinforcement learning algorithm for post-training LLMs. Unlike critic-based methods, GRPO computes advantages

Knowledge Beyond Language: Bridging the Gap in Multilingual Machine Unlearning Evaluation

ResearchDGX agent

arXiv:2605.14404v1 Announce Type: new Abstract: While LLMs are increasingly used in commercial services, they pose privacy risks such as leakage of sensitive personally identifiable information (PII).

Language Generation as Optimal Control: Closed-Loop Diffusion in Latent Control Space

SafetyDGX agent

arXiv:2605.14531v1 Announce Type: new Abstract: This work reformulates language generation as a stochastic optimal control problem, providing a unified theoretical perspective to analyze autoregressiv

Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards

SafetyDGX agent

arXiv:2605.14539v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an effective paradigm for improving the reasoning capabilities of large language mo

Leveraging Speech to Identify Signatures of Insight and Transfer in Problem Solving

ResearchDGX agent

arXiv:2605.12970v2 Announce Type: replace Abstract: Many problems seem to require a flash of insight to solve. What form do these sudden insights take, and what impact do they have on how people appro

LiSA: Lifelong Safety Adaptation via Conservative Policy Induction

Local AiDGX agent

arXiv:2605.14454v1 Announce Type: cross Abstract: As AI agents move from chat interfaces to systems that read private data, call tools, and execute multi-step workflows, guardrails become a last line

LLM-based Detection of Manipulative Political Narratives

ResearchDGX agent

arXiv:2605.14354v1 Announce Type: new Abstract: We present a new computational framework for detecting and structuring manipulative political narratives. A task that became more important due to the s

Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study

SafetyDGX agent

arXiv:2605.14087v1 Announce Type: new Abstract: Large Language Models (LLMs), when trained on web-scale corpora, inherently absorb toxic patterns from their training data. This leads to ``toxic degene

MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory

Model ReleasesDGX agent

arXiv:2605.15128v1 Announce Type: cross Abstract: Long-term agent memory is increasingly multimodal, yet existing evaluations rarely test whether agents preserve the visual evidence needed for later r

MemReranker: Reasoning-Aware Reranking for Agent Memory Retrieval

Model ReleasesDGX agent

arXiv:2605.06132v2 Announce Type: replace Abstract: In agent memory systems, the reranking model serves as the critical bridge connecting user queries with long-term memory. Most systems adopt the 're

Merging Methods for Multilingual Knowledge Editing for Large Language Models: An Empirical Odyssey

Model ReleasesDGX agent

arXiv:2605.13919v1 Announce Type: new Abstract: Multilingual knowledge editing (MKE) remains challenging because language-specific edits interfere with one another, even when locate-then-edit methods

MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs

SafetyDGX agent

arXiv:2605.15172v1 Announce Type: cross Abstract: Backdoor attacks pose a serious security threat to large language models (LLMs), which are increasingly deployed as general-purpose assistants in safe

Mini-JEPA Foundation Model Fleet Enables Agentic Hydrologic Intelligence

Model ReleasesDGX agent

arXiv:2605.14120v1 Announce Type: cross Abstract: Geospatial foundation models compress multispectral observations into dense embeddings increasingly used in natural-language environmental reasoning s

Minimal-Intervention KV Retention: A Design-Space Study and a Diversity-Penalty Survivor

Model ReleasesDGX agent

arXiv:2605.14292v1 Announce Type: cross Abstract: KV-cache compression at small budgets is a crowded design space spanning cache representation, head-wise routing, compression cadence, decoding behavi

Mining Subscenario Refactoring Opportunities in Behaviour-Driven Software Test Suites: ML Classifiers and LLM-Judge Baselines

Model ReleasesDGX agent

arXiv:2605.14568v1 Announce Type: cross Abstract: Context. Behaviour-Driven Development (BDD) software test suites accumulate duplicated step subsequences. Three published refactoring patterns are ava

Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding

Local AiDGX agent

arXiv:2605.14005v1 Announce Type: new Abstract: Speculative decoding has become a widely adopted technique for accelerating large language model (LLM) inference by drafting multiple candidate tokens a

Mitigating Data Scarcity in Psychological Defense Classification with Context-Aware Synthetic Augmentation

ResearchDGX agent

arXiv:2605.14380v1 Announce Type: new Abstract: Psychological defense mechanisms (PDMs) are unconscious cognitive processes that modulate how individuals perceive and respond to emotional distress. Au

MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring

Model ReleasesDGX agent

arXiv:2510.23477v2 Announce Type: replace Abstract: Effective math tutoring requires not only solving problems but also diagnosing students' difficulties and guiding them step by step. While multimoda

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning

SafetyDGX agent

arXiv:2512.07461v3 Announce Type: replace Abstract: We introduce Native Parallel Reasoner (NPR), a teacher-free framework that enables Large Language Models (LLMs) to self-evolve genuine parallel reas

Near-Miss: Latent Policy Failure Detection in Agentic Workflows

Model ReleasesDGX agent

arXiv:2603.29665v2 Announce Type: replace Abstract: Agentic systems for business process automation often require compliance with policies governing conditional updates to the system state. Evaluation

NodeSynth: Socially Aligned Synthetic Data for AI Evaluation

Model ReleasesDGX agent

arXiv:2605.14381v1 Announce Type: cross Abstract: Recent advancements in generative AI facilitate large-scale synthetic data generation for model evaluation. However, without targeted approaches, thes

Pause and Reflect: Conformal Aggregation for Chain-of-Thought Reasoning

ResearchDGX agent

arXiv:2605.14098v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning with self-consistency improves performance by aggregating multiple sampled reasoning paths. In this setting, correctn

Performance-Driven Policy Optimization for Speculative Decoding with Adaptive Windowing

SafetyDGX agent

arXiv:2605.14978v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by having a lightweight draft model propose speculative windows of candidate tokens for parallel verifica

Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music

SafetyDGX agent

arXiv:2605.14765v1 Announce Type: cross Abstract: Persian music, with its unique tonalities, modal systems (Dastgah), and rhythmic structures, presents significant challenges for music generation mode

Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning

Model ReleasesDGX agent

arXiv:2605.14040v1 Announce Type: new Abstract: We audit the multimodal-physics evaluation pipeline end-to-end and document three undetected construction practices that distort how the field measures

Polar probe linearly decodes semantic structures from LLMs

TutorialsDGX agent

arXiv:2605.14125v1 Announce Type: new Abstract: How do artificial neural networks bind concepts to form complex semantic structures? Here, we propose a simple neural code, whereby the existence and th

Proposal and study of statistical features for string similarity computation and classification

ResearchDGX agent

arXiv:2605.15110v1 Announce Type: cross Abstract: Adaptations of features commonly applied in the field of visual computing, co-occurrence matrix (COM) and run-length matrix (RLM), are proposed for th

Proxy Compression for Language Modeling

SafetyDGX agent

arXiv:2602.04289v2 Announce Type: replace Abstract: Modern language models are trained almost exclusively on token sequences produced by a fixed tokenizer, an external lossless compressor often over U

QOuLiPo: What a quantum computer sees when it reads a book

Model ReleasesDGX agent

arXiv:2605.14188v1 Announce Type: cross Abstract: What does a book look like to a quantum computer? This paper takes eight classical works of the Renaissance and its late-antique inheritance -- from A

Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases

SafetyDGX agent

arXiv:2601.03630v2 Announce Type: replace Abstract: This paper presents the first systematic comparison investigating whether Large Reasoning Models (LRMs) are superior judges to non-reasoning LLMs. O

Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax

SafetyDGX agent

arXiv:2605.14366v1 Announce Type: new Abstract: Extending large language models (LLMs) to low-resource languages often incurs an 'alignment tax': improvements in the target language come at the cost o

Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation

AgentsDGX agent

arXiv:2605.14563v1 Announce Type: cross Abstract: Automated code documentation is essential for modern software development, providing the contextual grounding that both human developers and coding ag

Rethinking Layer Relevance in Large Language Models Beyond Cosine Similarity

ResearchDGX agent

arXiv:2605.14075v1 Announce Type: cross Abstract: Large language models (LLMs) have revolutionized natural language processing. Understanding their internal mechanisms is crucial for developing more i

SciPaths: Forecasting Pathways to Scientific Discovery

Model ReleasesDGX agent

arXiv:2605.14600v1 Announce Type: new Abstract: Scientific progress depends on sequences of enabling contributions, yet existing AI4Science benchmarks largely focus on citation prediction, literature

Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility

Local AiDGX agent

arXiv:2605.14037v1 Announce Type: cross Abstract: Under modern test-time compute and agentic paradigms, language models process ever-longer sequences. Efficient text generation with transformer archit

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026)

Model ReleasesDGX agent

arXiv:2501.05465v2 Announce Type: replace Abstract: As foundation AI models continue to increase in size, an important question arises - is massive scale the only path forward? This survey of about 16

Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks

Model ReleasesDGX agent

arXiv:2605.15118v1 Announce Type: cross Abstract: We introduce a reusable framework for auditing whether LLM attack benchmarks collectively cover the threat surface: a 4imes6 Target imes Technique mat

The Scientific Contribution Graph: Automated Literature-based Technological Roadmapping at Scale

ResearchDGX agent

arXiv:2605.15011v1 Announce Type: new Abstract: Scientific contributions rarely develop in isolation, but instead build upon prior discoveries. We formulate the task of automated technological roadmap

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture

Model ReleasesDGX agent

arXiv:2605.14448v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have emerged as a powerful backbone for multimodal embeddings. Recent methods introduce chain-of-thought (CoT

Tokenizer Fertility and Zero-Shot Performance of Foundation Models on Ukrainian Legal Text: A Comparative Study

Model ReleasesDGX agent

arXiv:2605.14890v1 Announce Type: new Abstract: Foundation models tokenize Ukrainian legal text with vastly different efficiency, yet no systematic comparison exists for this domain. We benchmark seve

TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning

ResearchDGX agent

arXiv:2510.07118v3 Announce Type: replace Abstract: Instruction tuning is essential for aligning large language models (LLMs) to downstream tasks and commonly relies on large, diverse corpora. However

← Previous
1…7374757677…129
Next →