AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Model Releases

MAGA-Bench: Machine-Augment-Generated Text via Alignment Detection Benchmark

DGX agent

arXiv:2601.04633v2 Announce Type: replace Abstract: Machine-Generated Text (MGT) is becoming increasingly difficult to distinguish from Human-Written Text (HWT). This trend has exacerbated malicious a

model-releasesarxiv-cs-cl
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Mask the Target: A Plug-and-Play Regularizer Against LoRA Forgetting

DGX agent

arXiv:2605.29498v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become one of the most widely used fine-tuning mechanisms for adapting large language models to new domains, tasks, and u

safetyarxiv-cs-cl
29 May 2026
Model Releases

MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models

DGX agent

arXiv:2605.28825v1 Announce Type: new Abstract: Large language models (LLMs) frequently encode factual and reasoning knowledge in their internal representations that is not faithfully reflected in the

model-releasesarxiv-cs-cl
29 May 2026
Research

MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables

DGX agent

arXiv:2605.29859v1 Announce Type: cross Abstract: Recent speech language models rely on encoders that are optimized separately from autoregressive models. Since these encoders are unaware of the downs

researcharxiv-cs-cl
29 May 2026
Safety

Metric-Dependent Annotation Saturation for Learning from Label Distributions

DGX agent

arXiv:2605.29797v1 Announce Type: new Abstract: When annotators disagree on a label, the disagreement itself carries signal -- and the number of annotators needed to capture it depends on the evaluati

safetyarxiv-cs-cl
29 May 2026
Safety

MIC: Maximizing Informational Capacity in Adaptive Representations via Isotropic Subspace Alignment

DGX agent

arXiv:2605.29987v1 Announce Type: cross Abstract: Although multi-scales representation learning enables elastic-dimension embeddings, nested subspaces often suffer from dimensional redundancy and spec

safetyarxiv-cs-cl
29 May 2026
Research

Mindscape-Aware Retrieval Augmented Generation for Improved Long Context Understanding

DGX agent

arXiv:2512.17220v2 Announce Type: replace Abstract: Humans understand long and complex texts by relying on a holistic semantic representation of the content. This global view helps organize prior know

researcharxiv-cs-cl
29 May 2026
Model Releases

Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs

DGX agent

arXiv:2605.29737v1 Announce Type: cross Abstract: LLM-based coding assistants are seeing rapid adoption, offering substantial gains in developer productivity. As organizations increasingly ship code t

model-releasesarxiv-cs-cl
29 May 2026
Safety

Mining or Synthesis? Rethinking Exploration Efficiency in Iterative Alignment of Mathematical Reasoning

DGX agent

arXiv:2602.05370v3 Announce Type: replace Abstract: Iterative Direct Preference Optimization (DPO) has emerged as a widely used paradigm for aligning Large Language Models on reasoning tasks. Existing

safetyarxiv-cs-cl
29 May 2026
Research

Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context

DGX agent

arXiv:2510.06182v2 Announce Type: replace Abstract: A key component of in-context reasoning is the ability of language models (LMs) to bind entities for later retrieval. For example, an LM might repre

researcharxiv-cs-cl
29 May 2026
Research

Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels

DGX agent

arXiv:2605.29800v1 Announce Type: new Abstract: LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a frame

researcharxiv-cs-cl
29 May 2026
Agents

Offloading Score: Measuring AI Reliance Through Counterfactual Workflows

DGX agent

arXiv:2605.29392v1 Announce Type: cross Abstract: AI tools are increasingly integrated into real-world workflows. However, existing measures of reliance on these tools focus on AI output adoption or o

agentsarxiv-cs-cl
29 May 2026
Research

On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training

DGX agent

arXiv:2605.29496v1 Announce Type: new Abstract: Post-training has greatly improved reasoning in frontier vision-language models, yet its gains for perception remain comparatively limited, creating a b

researcharxiv-cs-cl
29 May 2026
Agents

PatchBoard: Schema-Grounded State Mutation for Reliable and Auditable LLM Multi-Agent Collaboration

DGX agent

arXiv:2605.29313v1 Announce Type: new Abstract: LLM multi-agent systems often coordinate through natural-language dialogue or loosely structured shared memory, making intermediate state difficult to v

agentsarxiv-cs-cl
29 May 2026
Safety

PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

DGX agent

arXiv:2605.29582v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown promise as educational tutors, yet effective tutoring requires more than solving problems: it must provide pro

safetyarxiv-cs-cl
29 May 2026
Research

Procedural Pretraining: Warming Up Language Models with Abstract Data

DGX agent

arXiv:2601.21725v2 Announce Type: replace Abstract: Pretraining language models directly on web-scale corpora is the de facto paradigm. We study an alternative where the model is initially exposed to

researcharxiv-cs-cl
29 May 2026
Local Ai

Prompt-Level Reward Specifications for Open-Ended Post-Training

DGX agent

arXiv:2605.29275v1 Announce Type: new Abstract: Open-ended post-training benefits from rewards that make prompt-specific success conditions explicit, rather than relying only on post-hoc scalar scores

local-aiarxiv-cs-cl
29 May 2026
Model Releases

Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization

DGX agent

arXiv:2605.29327v1 Announce Type: new Abstract: Efficient Distillation (EDistill) compresses large language models (LLMs) by structured pruning parameters and tuning lightweight modules with high trai

model-releasesarxiv-cs-cl
29 May 2026
Tutorials

Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models

DGX agent

arXiv:2605.28913v1 Announce Type: new Abstract: Large reasoning models (LRMs) often generate extensive chain-of-thought (CoT) traces before producing a final answer. As explicit textual artifacts, the

tutorialsarxiv-cs-cl
29 May 2026
Model Releases

Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs

DGX agent

arXiv:2605.30021v1 Announce Type: new Abstract: Many open-ended instructions have multiple valid answers that users can benefit from seeing, but post-training often narrows an LLM's output space towar

model-releasesarxiv-cs-cl
29 May 2026
Safety

Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents

DGX agent

arXiv:2605.29447v1 Announce Type: cross Abstract: While GUI agents have advanced rapidly, they often lack the robustness to recover from their own errors, hindering real-world deployment. To bridge th

safetyarxiv-cs-cl
29 May 2026
Research

Resolution Diagnostics for Paired LLM Evaluation

DGX agent

arXiv:2605.30315v1 Announce Type: new Abstract: Across two public LLM leaderboards, many displayed pairwise rankings do not meet a conventional paired-test resolution target under the actual paired ev

researcharxiv-cs-cl
29 May 2026
Research

Rethinking Stepwise Model Routing: A Cost-Efficient Table Reasoning Perspective

DGX agent

arXiv:2605.29319v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on table reasoning tasks but incur substantial inference cost due to long reasoning traces. Ste

researcharxiv-cs-cl
29 May 2026
Agents

Revisiting Observation Reduction for Web Agents: Comprehensive Evaluation with a Lightweight Framework

DGX agent

arXiv:2605.29397v1 Announce Type: new Abstract: HTML observations in LLM-based web agents are extremely long, and while many reduction methods have been proposed, it remains unclear which methods redu

agentsarxiv-cs-cl
29 May 2026
Model Releases

RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment

DGX agent

arXiv:2605.28827v1 Announce Type: new Abstract: Open Arabic large language models split into two classes: sub-1B multilingual models that treat Arabic as an afterthought (Qwen2.5-0.5B, Falcon-H1-0.5B)

model-releasesarxiv-cs-cl
29 May 2026
Safety

RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains

DGX agent

arXiv:2605.29156v1 Announce Type: cross Abstract: Pointwise reward modeling offers critical signals for LLM post-training, yet struggles with absolute scoring in subjective, non-verifiable settings. R

safetyarxiv-cs-cl
29 May 2026
Model Releases

Scaling Laws for Agent Harnesses via Effective Feedback Compute

DGX agent

arXiv:2605.29682v1 Announce Type: new Abstract: Agent harnesses increasingly determine the performance of language-model systems by deciding how models call tools, receive feedback, verify intermediat

model-releasesarxiv-cs-cl
29 May 2026
Agents

SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?

DGX agent

arXiv:2605.30104v1 Announce Type: new Abstract: Widely used language-model benchmarks are increasingly saturated, with frontier systems often receiving near-tied scores that standard metrics cannot re

agentsarxiv-cs-cl
29 May 2026
Research

ShapleyLaw: A Game-Theoretic Approach to Multilingual Scaling Laws

DGX agent

arXiv:2603.17945v2 Announce Type: replace Abstract: In multilingual pretraining, the test loss of a pretrained model is heavily influenced by the proportion of each language in the pretraining data, n

researcharxiv-cs-cl
29 May 2026
Research

Slogans or Stance? A Label-Light Diagnostic for Entrepreneurial-Discourse Measurement on Chinese SOE Speeches

DGX agent

arXiv:2605.29188v1 Announce Type: new Abstract: Dictionary methods, topic models, and embedding-similarity scorers are widely used in CSS and management research to measure constructs such as 'entrepr

researcharxiv-cs-cl
29 May 2026
Research

Spurious Prompts: Can Irrelevant Prompts Steer Large Language Models?

DGX agent

arXiv:2605.29678v1 Announce Type: new Abstract: Large language models are highly sensitive to prompts, but this sensitivity is usually studied through task-relevant instructions, demonstrations, or re

researcharxiv-cs-cl
29 May 2026
Model Releases

STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments

DGX agent

arXiv:2605.29324v1 Announce Type: new Abstract: Mobile GUI agents excel at immediate reactive control but frequently fail in realistic, long-horizon tasks that require memory. This failure stems from

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Text-Preserving Lossy Text Compression: A Study of Strategic Deletion and LLM Reconstruction

DGX agent

arXiv:2605.29000v1 Announce Type: new Abstract: Traditional lossless text compression preserves every byte, but its gains on natural language are often modest in realistic operating regimes. We study

model-releasesarxiv-cs-cl
29 May 2026
Safety

The Anatomy of Conversational Scams: A Topic-Based Red Teaming Analysis of Multi-Turn Interactions in LLMs

DGX agent

arXiv:2601.03134v2 Announce Type: replace Abstract: As LLMs gain persuasive capabilities through extended dialogues, they create new opportunities for studying adversarial conversational behavior in e

safetyarxiv-cs-cl
29 May 2026
Model Releases

The Trust Paradox: How CS Researchers Engage LLM Leaderboards

DGX agent

arXiv:2605.28966v1 Announce Type: new Abstract: Large language model (LLM) leaderboards rank AI models using standardized benchmarks and have become highly visible across computer science, despite kno

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

DGX agent

arXiv:2602.15382v2 Announce Type: replace Abstract: Multi-Agent Systems (MAS) powered by Large Language Models have unlocked advanced collaborative reasoning, yet they remain bottlenecked by discrete

model-releasesarxiv-cs-cl
29 May 2026
Research

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge

DGX agent

arXiv:2505.16178v2 Announce Type: replace Abstract: While fine-tuning is the standard for injecting factual knowledge into large language models (LLMs), the mechanisms enabling reliable fact recall vi

researcharxiv-cs-cl
29 May 2026
Model Releases

Understanding Safety-Sensitive Expert Behavior in Mixture-of-Experts LLMs

DGX agent

arXiv:2605.29708v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) LLMs rely on sparse, router-driven expert activation, yet how safety alignment interacts with routed expert specialization rema

model-releasesarxiv-cs-cl
29 May 2026
Research

Understanding the Ability of LLMs to Handle Character-Level Perturbation

DGX agent

arXiv:2510.14365v4 Announce Type: replace Abstract: This work investigates the resilience of contemporary large language models (LLMs) against frequent character-level perturbations. We examine three

researcharxiv-cs-cl
29 May 2026
Research

UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering

DGX agent

arXiv:2605.30076v1 Announce Type: new Abstract: Activation-based control steers large language models (LLMs) by intervening on their internal representations during inference, and has emerged as an ef

researcharxiv-cs-cl
29 May 2026
Safety

User-Aware Active Knowledge Acquisition for Emotional Support Dialogue

DGX agent

arXiv:2605.29715v1 Announce Type: new Abstract: Emotional support plays an important role in dialogue systems, and its success depends on adapting to a user's evolving and implicit needs across multi-

safetyarxiv-cs-cl
29 May 2026
Research

Valency Classification of Mapudungun Verbal Roots. Established by the language's own morphotactics

DGX agent

arXiv:2604.00789v3 Announce Type: replace Abstract: In the previous work, a lexical (re)categorisation -- or confirmation of the given category -- of roots identified as verbal was undertaken to deter

researcharxiv-cs-cl
29 May 2026
Safety

ValueFlow: Measuring the Propagation of Value Perturbations in Multi-Agent LLM Systems

DGX agent

arXiv:2602.08567v2 Announce Type: replace-cross Abstract: Multi-agent large language model (LLM) systems increasingly consist of agents that observe and respond to one another's outputs. While value a

safetyarxiv-cs-cl
29 May 2026
Model Releases

Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering

DGX agent

arXiv:2605.29648v1 Announce Type: new Abstract: Applying reinforcement learning to improve factual accuracy in knowledge-intensive question answering faces a reward design dilemma. Response-level rewa

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents

DGX agent

arXiv:2605.30256v1 Announce Type: cross Abstract: Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonve

model-releasesarxiv-cs-cl
29 May 2026
Research

WaterSearch: A Quality-Aware Search-based Watermarking Framework for Large Language Models

DGX agent

arXiv:2512.00837v2 Announce Type: replace Abstract: Watermarking acts as a critical safeguard in text generated by Large Language Models (LLMs). By embedding identifiable signals into model outputs, w

researcharxiv-cs-cl
29 May 2026
Research

What are They Thinking? Delineation, Probing and Tracking of Concepts in LLMs

DGX agent

arXiv:2605.28823v1 Announce Type: new Abstract: As the influence of LLMs expands, it is imperative to gain insight into their decisions. One way to do that is to develop probes that detect the presenc

researcharxiv-cs-cl
29 May 2026
Tutorials

What Exactly do Children Receive in Language Acquisition? A Case Study on CHILDES with Automated Detection of Filler-Gap Dependencies

DGX agent

arXiv:2603.02082v2 Announce Type: replace Abstract: Children's acquisition of filler-gap dependencies has been argued by some to depend on innate grammatical knowledge, while others suggest that the d

tutorialsarxiv-cs-cl
29 May 2026
← Previous
1…6970717273…162
Next →