AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlog
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Model Releases

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

DGX agent

arXiv:2605.15104v1 Announce Type: new Abstract: Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified

model-releasesarxiv-cs-cl
15 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

Geometry-Aware Decoding with Wasserstein-Regularized Truncation and Mass Penalties for Large Language Models

DGX agent

arXiv:2602.10346v2 Announce Type: replace Abstract: Large language models (LLMs) must balance diversity and creativity against logical coherence in open-ended generation. Existing truncation-based sam

researcharxiv-cs-cl
15 May 2026
Safety

GradShield: Alignment Preserving Finetuning

DGX agent

arXiv:2605.14194v1 Announce Type: new Abstract: Large Language Models (LLMs) pose a significant risk of safety misalignment after finetuning, as models can be compromised by both explicitly and implic

safetyarxiv-cs-cl
15 May 2026
Model Releases

GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations

DGX agent

arXiv:2605.14498v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly serve as personal assistants and workplace collaborators, where their utility depends on memory systems t

model-releasesarxiv-cs-cl
15 May 2026
Safety

Ideology Prediction of German Political Texts

DGX agent

arXiv:2605.14352v1 Announce Type: new Abstract: Elections represent a crucial milestone in a nation's ongoing development. To better understand the political rhetoric from various movements, ranging f

safetyarxiv-cs-cl
15 May 2026
Model Releases

Is Grep All You Need? How Agent Harnesses Reshape Agentic Search

DGX agent

arXiv:2605.15184v1 Announce Type: new Abstract: Recent advances in Large Language Model (LLM) agents have enabled complex agentic workflows where models autonomously retrieve information, call tools,

model-releasesarxiv-cs-cl
15 May 2026
Research

It Takes Two: Your GRPO Is Secretly DPO

DGX agent

arXiv:2510.00977v3 Announce Type: replace-cross Abstract: GRPO has emerged as a prominent reinforcement learning algorithm for post-training LLMs. Unlike critic-based methods, GRPO computes advantages

researcharxiv-cs-cl
15 May 2026
Research

Knowledge Beyond Language: Bridging the Gap in Multilingual Machine Unlearning Evaluation

DGX agent

arXiv:2605.14404v1 Announce Type: new Abstract: While LLMs are increasingly used in commercial services, they pose privacy risks such as leakage of sensitive personally identifiable information (PII).

researcharxiv-cs-cl
15 May 2026
Safety

Language Generation as Optimal Control: Closed-Loop Diffusion in Latent Control Space

DGX agent

arXiv:2605.14531v1 Announce Type: new Abstract: This work reformulates language generation as a stochastic optimal control problem, providing a unified theoretical perspective to analyze autoregressiv

safetyarxiv-cs-cl
15 May 2026
Safety

Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards

DGX agent

arXiv:2605.14539v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an effective paradigm for improving the reasoning capabilities of large language mo

safetyarxiv-cs-cl
15 May 2026
Research

Leveraging Speech to Identify Signatures of Insight and Transfer in Problem Solving

DGX agent

arXiv:2605.12970v2 Announce Type: replace Abstract: Many problems seem to require a flash of insight to solve. What form do these sudden insights take, and what impact do they have on how people appro

researcharxiv-cs-cl
15 May 2026
Local Ai

LiSA: Lifelong Safety Adaptation via Conservative Policy Induction

DGX agent

arXiv:2605.14454v1 Announce Type: cross Abstract: As AI agents move from chat interfaces to systems that read private data, call tools, and execute multi-step workflows, guardrails become a last line

local-aiarxiv-cs-cl
15 May 2026
Research

LLM-based Detection of Manipulative Political Narratives

DGX agent

arXiv:2605.14354v1 Announce Type: new Abstract: We present a new computational framework for detecting and structuring manipulative political narratives. A task that became more important due to the s

researcharxiv-cs-cl
15 May 2026
Safety

Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study

DGX agent

arXiv:2605.14087v1 Announce Type: new Abstract: Large Language Models (LLMs), when trained on web-scale corpora, inherently absorb toxic patterns from their training data. This leads to ``toxic degene

safetyarxiv-cs-cl
15 May 2026
Model Releases

MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory

DGX agent

arXiv:2605.15128v1 Announce Type: cross Abstract: Long-term agent memory is increasingly multimodal, yet existing evaluations rarely test whether agents preserve the visual evidence needed for later r

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

MemReranker: Reasoning-Aware Reranking for Agent Memory Retrieval

DGX agent

arXiv:2605.06132v2 Announce Type: replace Abstract: In agent memory systems, the reranking model serves as the critical bridge connecting user queries with long-term memory. Most systems adopt the 're

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

Merging Methods for Multilingual Knowledge Editing for Large Language Models: An Empirical Odyssey

DGX agent

arXiv:2605.13919v1 Announce Type: new Abstract: Multilingual knowledge editing (MKE) remains challenging because language-specific edits interfere with one another, even when locate-then-edit methods

model-releasesarxiv-cs-cl
15 May 2026
Safety

MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs

DGX agent

arXiv:2605.15172v1 Announce Type: cross Abstract: Backdoor attacks pose a serious security threat to large language models (LLMs), which are increasingly deployed as general-purpose assistants in safe

safetyarxiv-cs-cl
15 May 2026
Model Releases

Mini-JEPA Foundation Model Fleet Enables Agentic Hydrologic Intelligence

DGX agent

arXiv:2605.14120v1 Announce Type: cross Abstract: Geospatial foundation models compress multispectral observations into dense embeddings increasingly used in natural-language environmental reasoning s

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

Minimal-Intervention KV Retention: A Design-Space Study and a Diversity-Penalty Survivor

DGX agent

arXiv:2605.14292v1 Announce Type: cross Abstract: KV-cache compression at small budgets is a crowded design space spanning cache representation, head-wise routing, compression cadence, decoding behavi

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

Mining Subscenario Refactoring Opportunities in Behaviour-Driven Software Test Suites: ML Classifiers and LLM-Judge Baselines

DGX agent

arXiv:2605.14568v1 Announce Type: cross Abstract: Context. Behaviour-Driven Development (BDD) software test suites accumulate duplicated step subsequences. Three published refactoring patterns are ava

model-releasesarxiv-cs-cl
15 May 2026
Local Ai

Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding

DGX agent

arXiv:2605.14005v1 Announce Type: new Abstract: Speculative decoding has become a widely adopted technique for accelerating large language model (LLM) inference by drafting multiple candidate tokens a

local-aiarxiv-cs-cl
15 May 2026
Research

Mitigating Data Scarcity in Psychological Defense Classification with Context-Aware Synthetic Augmentation

DGX agent

arXiv:2605.14380v1 Announce Type: new Abstract: Psychological defense mechanisms (PDMs) are unconscious cognitive processes that modulate how individuals perceive and respond to emotional distress. Au

researcharxiv-cs-cl
15 May 2026
Model Releases

MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring

DGX agent

arXiv:2510.23477v2 Announce Type: replace Abstract: Effective math tutoring requires not only solving problems but also diagnosing students' difficulties and guiding them step by step. While multimoda

model-releasesarxiv-cs-cl
15 May 2026
Safety

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning

DGX agent

arXiv:2512.07461v3 Announce Type: replace Abstract: We introduce Native Parallel Reasoner (NPR), a teacher-free framework that enables Large Language Models (LLMs) to self-evolve genuine parallel reas

safetyarxiv-cs-cl
15 May 2026
Model Releases

Near-Miss: Latent Policy Failure Detection in Agentic Workflows

DGX agent

arXiv:2603.29665v2 Announce Type: replace Abstract: Agentic systems for business process automation often require compliance with policies governing conditional updates to the system state. Evaluation

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

NodeSynth: Socially Aligned Synthetic Data for AI Evaluation

DGX agent

arXiv:2605.14381v1 Announce Type: cross Abstract: Recent advancements in generative AI facilitate large-scale synthetic data generation for model evaluation. However, without targeted approaches, thes

model-releasesarxiv-cs-cl
15 May 2026
Research

Pause and Reflect: Conformal Aggregation for Chain-of-Thought Reasoning

DGX agent

arXiv:2605.14098v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning with self-consistency improves performance by aggregating multiple sampled reasoning paths. In this setting, correctn

researcharxiv-cs-cl
15 May 2026
Safety

Performance-Driven Policy Optimization for Speculative Decoding with Adaptive Windowing

DGX agent

arXiv:2605.14978v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by having a lightweight draft model propose speculative windows of candidate tokens for parallel verifica

safetyarxiv-cs-cl
15 May 2026
Safety

Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music

DGX agent

arXiv:2605.14765v1 Announce Type: cross Abstract: Persian music, with its unique tonalities, modal systems (Dastgah), and rhythmic structures, presents significant challenges for music generation mode

safetyarxiv-cs-cl
15 May 2026
Model Releases

Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning

DGX agent

arXiv:2605.14040v1 Announce Type: new Abstract: We audit the multimodal-physics evaluation pipeline end-to-end and document three undetected construction practices that distort how the field measures

model-releasesarxiv-cs-cl
15 May 2026
Tutorials

Polar probe linearly decodes semantic structures from LLMs

DGX agent

arXiv:2605.14125v1 Announce Type: new Abstract: How do artificial neural networks bind concepts to form complex semantic structures? Here, we propose a simple neural code, whereby the existence and th

tutorialsarxiv-cs-cl
15 May 2026
Research

Proposal and study of statistical features for string similarity computation and classification

DGX agent

arXiv:2605.15110v1 Announce Type: cross Abstract: Adaptations of features commonly applied in the field of visual computing, co-occurrence matrix (COM) and run-length matrix (RLM), are proposed for th

researcharxiv-cs-cl
15 May 2026
Safety

Proxy Compression for Language Modeling

DGX agent

arXiv:2602.04289v2 Announce Type: replace Abstract: Modern language models are trained almost exclusively on token sequences produced by a fixed tokenizer, an external lossless compressor often over U

safetyarxiv-cs-cl
15 May 2026
Model Releases

QOuLiPo: What a quantum computer sees when it reads a book

DGX agent

arXiv:2605.14188v1 Announce Type: cross Abstract: What does a book look like to a quantum computer? This paper takes eight classical works of the Renaissance and its late-antique inheritance -- from A

model-releasesarxiv-cs-cl
15 May 2026
Safety

Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases

DGX agent

arXiv:2601.03630v2 Announce Type: replace Abstract: This paper presents the first systematic comparison investigating whether Large Reasoning Models (LRMs) are superior judges to non-reasoning LLMs. O

safetyarxiv-cs-cl
15 May 2026
Safety

Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax

DGX agent

arXiv:2605.14366v1 Announce Type: new Abstract: Extending large language models (LLMs) to low-resource languages often incurs an 'alignment tax': improvements in the target language come at the cost o

safetyarxiv-cs-cl
15 May 2026
Agents

Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation

DGX agent

arXiv:2605.14563v1 Announce Type: cross Abstract: Automated code documentation is essential for modern software development, providing the contextual grounding that both human developers and coding ag

agentsarxiv-cs-cl
15 May 2026
Research

Rethinking Layer Relevance in Large Language Models Beyond Cosine Similarity

DGX agent

arXiv:2605.14075v1 Announce Type: cross Abstract: Large language models (LLMs) have revolutionized natural language processing. Understanding their internal mechanisms is crucial for developing more i

researcharxiv-cs-cl
15 May 2026
Model Releases

SciPaths: Forecasting Pathways to Scientific Discovery

DGX agent

arXiv:2605.14600v1 Announce Type: new Abstract: Scientific progress depends on sequences of enabling contributions, yet existing AI4Science benchmarks largely focus on citation prediction, literature

model-releasesarxiv-cs-cl
15 May 2026
Local Ai

Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility

DGX agent

arXiv:2605.14037v1 Announce Type: cross Abstract: Under modern test-time compute and agentic paradigms, language models process ever-longer sequences. Efficient text generation with transformer archit

local-aiarxiv-cs-cl
15 May 2026
Model Releases

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026)

DGX agent

arXiv:2501.05465v2 Announce Type: replace Abstract: As foundation AI models continue to increase in size, an important question arises - is massive scale the only path forward? This survey of about 16

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks

DGX agent

arXiv:2605.15118v1 Announce Type: cross Abstract: We introduce a reusable framework for auditing whether LLM attack benchmarks collectively cover the threat surface: a 4imes6 Target imes Technique mat

model-releasesarxiv-cs-cl
15 May 2026
Research

The Scientific Contribution Graph: Automated Literature-based Technological Roadmapping at Scale

DGX agent

arXiv:2605.15011v1 Announce Type: new Abstract: Scientific contributions rarely develop in isolation, but instead build upon prior discoveries. We formulate the task of automated technological roadmap

researcharxiv-cs-cl
15 May 2026
Model Releases

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture

DGX agent

arXiv:2605.14448v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have emerged as a powerful backbone for multimodal embeddings. Recent methods introduce chain-of-thought (CoT

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

Tokenizer Fertility and Zero-Shot Performance of Foundation Models on Ukrainian Legal Text: A Comparative Study

DGX agent

arXiv:2605.14890v1 Announce Type: new Abstract: Foundation models tokenize Ukrainian legal text with vastly different efficiency, yet no systematic comparison exists for this domain. We benchmark seve

model-releasesarxiv-cs-cl
15 May 2026
Research

TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning

DGX agent

arXiv:2510.07118v3 Announce Type: replace Abstract: Instruction tuning is essential for aligning large language models (LLMs) to downstream tasks and commonly relies on large, diverse corpora. However

researcharxiv-cs-cl
15 May 2026
Research

Uncertainty Quantification for Large Language Diffusion Models

DGX agent

arXiv:2605.14570v1 Announce Type: new Abstract: Large Language Diffusion Models (LLDMs) are emerging as an alternative to autoregressive models, offering faster inference through higher parallelism. S

researcharxiv-cs-cl
15 May 2026
← Previous
1…9394959697…162
Next →