AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Agents

Offloading Score: Measuring AI Reliance Through Counterfactual Workflows

DGX agent

arXiv:2605.29392v1 Announce Type: cross Abstract: AI tools are increasingly integrated into real-world workflows. However, existing measures of reliance on these tools focus on AI output adoption or o

agentsarxiv-cs-cl
29 May 2026
Research
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training

DGX agent

arXiv:2605.29496v1 Announce Type: new Abstract: Post-training has greatly improved reasoning in frontier vision-language models, yet its gains for perception remain comparatively limited, creating a b

researcharxiv-cs-cl
29 May 2026
Agents

PatchBoard: Schema-Grounded State Mutation for Reliable and Auditable LLM Multi-Agent Collaboration

DGX agent

arXiv:2605.29313v1 Announce Type: new Abstract: LLM multi-agent systems often coordinate through natural-language dialogue or loosely structured shared memory, making intermediate state difficult to v

agentsarxiv-cs-cl
29 May 2026
Safety

PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

DGX agent

arXiv:2605.29582v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown promise as educational tutors, yet effective tutoring requires more than solving problems: it must provide pro

safetyarxiv-cs-cl
29 May 2026
Research

Procedural Pretraining: Warming Up Language Models with Abstract Data

DGX agent

arXiv:2601.21725v2 Announce Type: replace Abstract: Pretraining language models directly on web-scale corpora is the de facto paradigm. We study an alternative where the model is initially exposed to

researcharxiv-cs-cl
29 May 2026
Local Ai

Prompt-Level Reward Specifications for Open-Ended Post-Training

DGX agent

arXiv:2605.29275v1 Announce Type: new Abstract: Open-ended post-training benefits from rewards that make prompt-specific success conditions explicit, rather than relying only on post-hoc scalar scores

local-aiarxiv-cs-cl
29 May 2026
Model Releases

Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization

DGX agent

arXiv:2605.29327v1 Announce Type: new Abstract: Efficient Distillation (EDistill) compresses large language models (LLMs) by structured pruning parameters and tuning lightweight modules with high trai

model-releasesarxiv-cs-cl
29 May 2026
Tutorials

Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models

DGX agent

arXiv:2605.28913v1 Announce Type: new Abstract: Large reasoning models (LRMs) often generate extensive chain-of-thought (CoT) traces before producing a final answer. As explicit textual artifacts, the

tutorialsarxiv-cs-cl
29 May 2026
Model Releases

Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs

DGX agent

arXiv:2605.30021v1 Announce Type: new Abstract: Many open-ended instructions have multiple valid answers that users can benefit from seeing, but post-training often narrows an LLM's output space towar

model-releasesarxiv-cs-cl
29 May 2026
Safety

Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents

DGX agent

arXiv:2605.29447v1 Announce Type: cross Abstract: While GUI agents have advanced rapidly, they often lack the robustness to recover from their own errors, hindering real-world deployment. To bridge th

safetyarxiv-cs-cl
29 May 2026
Research

Resolution Diagnostics for Paired LLM Evaluation

DGX agent

arXiv:2605.30315v1 Announce Type: new Abstract: Across two public LLM leaderboards, many displayed pairwise rankings do not meet a conventional paired-test resolution target under the actual paired ev

researcharxiv-cs-cl
29 May 2026
Research

Rethinking Stepwise Model Routing: A Cost-Efficient Table Reasoning Perspective

DGX agent

arXiv:2605.29319v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on table reasoning tasks but incur substantial inference cost due to long reasoning traces. Ste

researcharxiv-cs-cl
29 May 2026
Agents

Revisiting Observation Reduction for Web Agents: Comprehensive Evaluation with a Lightweight Framework

DGX agent

arXiv:2605.29397v1 Announce Type: new Abstract: HTML observations in LLM-based web agents are extremely long, and while many reduction methods have been proposed, it remains unclear which methods redu

agentsarxiv-cs-cl
29 May 2026
Model Releases

RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment

DGX agent

arXiv:2605.28827v1 Announce Type: new Abstract: Open Arabic large language models split into two classes: sub-1B multilingual models that treat Arabic as an afterthought (Qwen2.5-0.5B, Falcon-H1-0.5B)

model-releasesarxiv-cs-cl
29 May 2026
Safety

RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains

DGX agent

arXiv:2605.29156v1 Announce Type: cross Abstract: Pointwise reward modeling offers critical signals for LLM post-training, yet struggles with absolute scoring in subjective, non-verifiable settings. R

safetyarxiv-cs-cl
29 May 2026
Model Releases

Scaling Laws for Agent Harnesses via Effective Feedback Compute

DGX agent

arXiv:2605.29682v1 Announce Type: new Abstract: Agent harnesses increasingly determine the performance of language-model systems by deciding how models call tools, receive feedback, verify intermediat

model-releasesarxiv-cs-cl
29 May 2026
Agents

SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?

DGX agent

arXiv:2605.30104v1 Announce Type: new Abstract: Widely used language-model benchmarks are increasingly saturated, with frontier systems often receiving near-tied scores that standard metrics cannot re

agentsarxiv-cs-cl
29 May 2026
Research

ShapleyLaw: A Game-Theoretic Approach to Multilingual Scaling Laws

DGX agent

arXiv:2603.17945v2 Announce Type: replace Abstract: In multilingual pretraining, the test loss of a pretrained model is heavily influenced by the proportion of each language in the pretraining data, n

researcharxiv-cs-cl
29 May 2026
Research

Slogans or Stance? A Label-Light Diagnostic for Entrepreneurial-Discourse Measurement on Chinese SOE Speeches

DGX agent

arXiv:2605.29188v1 Announce Type: new Abstract: Dictionary methods, topic models, and embedding-similarity scorers are widely used in CSS and management research to measure constructs such as 'entrepr

researcharxiv-cs-cl
29 May 2026
Research

Spurious Prompts: Can Irrelevant Prompts Steer Large Language Models?

DGX agent

arXiv:2605.29678v1 Announce Type: new Abstract: Large language models are highly sensitive to prompts, but this sensitivity is usually studied through task-relevant instructions, demonstrations, or re

researcharxiv-cs-cl
29 May 2026
Model Releases

STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments

DGX agent

arXiv:2605.29324v1 Announce Type: new Abstract: Mobile GUI agents excel at immediate reactive control but frequently fail in realistic, long-horizon tasks that require memory. This failure stems from

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Text-Preserving Lossy Text Compression: A Study of Strategic Deletion and LLM Reconstruction

DGX agent

arXiv:2605.29000v1 Announce Type: new Abstract: Traditional lossless text compression preserves every byte, but its gains on natural language are often modest in realistic operating regimes. We study

model-releasesarxiv-cs-cl
29 May 2026
Safety

The Anatomy of Conversational Scams: A Topic-Based Red Teaming Analysis of Multi-Turn Interactions in LLMs

DGX agent

arXiv:2601.03134v2 Announce Type: replace Abstract: As LLMs gain persuasive capabilities through extended dialogues, they create new opportunities for studying adversarial conversational behavior in e

safetyarxiv-cs-cl
29 May 2026
Model Releases

The Trust Paradox: How CS Researchers Engage LLM Leaderboards

DGX agent

arXiv:2605.28966v1 Announce Type: new Abstract: Large language model (LLM) leaderboards rank AI models using standardized benchmarks and have become highly visible across computer science, despite kno

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

DGX agent

arXiv:2602.15382v2 Announce Type: replace Abstract: Multi-Agent Systems (MAS) powered by Large Language Models have unlocked advanced collaborative reasoning, yet they remain bottlenecked by discrete

model-releasesarxiv-cs-cl
29 May 2026
Research

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge

DGX agent

arXiv:2505.16178v2 Announce Type: replace Abstract: While fine-tuning is the standard for injecting factual knowledge into large language models (LLMs), the mechanisms enabling reliable fact recall vi

researcharxiv-cs-cl
29 May 2026
Model Releases

Understanding Safety-Sensitive Expert Behavior in Mixture-of-Experts LLMs

DGX agent

arXiv:2605.29708v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) LLMs rely on sparse, router-driven expert activation, yet how safety alignment interacts with routed expert specialization rema

model-releasesarxiv-cs-cl
29 May 2026
Research

Understanding the Ability of LLMs to Handle Character-Level Perturbation

DGX agent

arXiv:2510.14365v4 Announce Type: replace Abstract: This work investigates the resilience of contemporary large language models (LLMs) against frequent character-level perturbations. We examine three

researcharxiv-cs-cl
29 May 2026
Research

UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering

DGX agent

arXiv:2605.30076v1 Announce Type: new Abstract: Activation-based control steers large language models (LLMs) by intervening on their internal representations during inference, and has emerged as an ef

researcharxiv-cs-cl
29 May 2026
Safety

User-Aware Active Knowledge Acquisition for Emotional Support Dialogue

DGX agent

arXiv:2605.29715v1 Announce Type: new Abstract: Emotional support plays an important role in dialogue systems, and its success depends on adapting to a user's evolving and implicit needs across multi-

safetyarxiv-cs-cl
29 May 2026
Research

Valency Classification of Mapudungun Verbal Roots. Established by the language's own morphotactics

DGX agent

arXiv:2604.00789v3 Announce Type: replace Abstract: In the previous work, a lexical (re)categorisation -- or confirmation of the given category -- of roots identified as verbal was undertaken to deter

researcharxiv-cs-cl
29 May 2026
Safety

ValueFlow: Measuring the Propagation of Value Perturbations in Multi-Agent LLM Systems

DGX agent

arXiv:2602.08567v2 Announce Type: replace-cross Abstract: Multi-agent large language model (LLM) systems increasingly consist of agents that observe and respond to one another's outputs. While value a

safetyarxiv-cs-cl
29 May 2026
Model Releases

Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering

DGX agent

arXiv:2605.29648v1 Announce Type: new Abstract: Applying reinforcement learning to improve factual accuracy in knowledge-intensive question answering faces a reward design dilemma. Response-level rewa

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents

DGX agent

arXiv:2605.30256v1 Announce Type: cross Abstract: Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonve

model-releasesarxiv-cs-cl
29 May 2026
Research

WaterSearch: A Quality-Aware Search-based Watermarking Framework for Large Language Models

DGX agent

arXiv:2512.00837v2 Announce Type: replace Abstract: Watermarking acts as a critical safeguard in text generated by Large Language Models (LLMs). By embedding identifiable signals into model outputs, w

researcharxiv-cs-cl
29 May 2026
Research

What are They Thinking? Delineation, Probing and Tracking of Concepts in LLMs

DGX agent

arXiv:2605.28823v1 Announce Type: new Abstract: As the influence of LLMs expands, it is imperative to gain insight into their decisions. One way to do that is to develop probes that detect the presenc

researcharxiv-cs-cl
29 May 2026
Tutorials

What Exactly do Children Receive in Language Acquisition? A Case Study on CHILDES with Automated Detection of Filler-Gap Dependencies

DGX agent

arXiv:2603.02082v2 Announce Type: replace Abstract: Children's acquisition of filler-gap dependencies has been argued by some to depend on innate grammatical knowledge, while others suggest that the d

tutorialsarxiv-cs-cl
29 May 2026
Research

When RL Suppresses Its Own Vocabulary: Recovering Reasoning Diversity in Puzzle-to-Math Transfer

DGX agent

arXiv:2605.29190v1 Announce Type: cross Abstract: Reinforcement learning using verifiable rewards (RLVR) improves LLM reasoning, but the conditions under which it transfers across domains -- and why i

researcharxiv-cs-cl
29 May 2026
Model Releases

When the Same Coefficients Reach Different Places: Asymmetric Realizability in Transplanting Tokenizers across Large Language Models

DGX agent

arXiv:2601.00065v3 Announce Type: replace-cross Abstract: Tokenizer transplant in cross-vocabulary model composition reconstructs donor-only embedding rows as weighted combinations over shared lexical

model-releasesarxiv-cs-cl
29 May 2026
Applications

Who Am I? History-Aware Profiles for Student Simulation in Tutoring Dialogues

DGX agent

arXiv:2605.30051v1 Announce Type: new Abstract: A key part of developing large language model (LLM)-powered, automated tutoring tools is student simulation, i.e., using LLMs to role-play as students,

applicationsarxiv-cs-cl
29 May 2026
Model Releases

World Models in Words: Auditing Physical State-Transition Commitments in Vision-Language Models

DGX agent

arXiv:2605.29585v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used to answer questions about physical scenes, yet most evaluations reduce performance to a final answer

model-releasesarxiv-cs-cl
29 May 2026
Agents

WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction

DGX agent

arXiv:2605.29341v1 Announce Type: cross Abstract: Multimodal large language models are increasingly deployed as long-horizon agents, where memory must do more than recall: it must track an evolving wo

agentsarxiv-cs-cl
29 May 2026
Research

X-GS: An Extensible Framework for Perceiving and Thinking via 3D Gaussian Splatting

DGX agent

arXiv:2603.09632v3 Announce Type: replace-cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful technique for novel view synthesis, subsequently extending into numerous spatial AI app

researcharxiv-cs-cl
29 May 2026
Research

A new semantically annotated corpus with syntactic-semantic and cross-lingual senses

DGX agent

arXiv:2605.28494v1 Announce Type: new Abstract: We describe a new sense-tagged corpus for word sense disambiguation. The corpus is constituted of instances of 20 French polysemous verbs. Each verb ins

researcharxiv-cs-cl
28 May 2026
Research

A tree interpretation of arc standard dependency derivation

DGX agent

arXiv:2603.27459v2 Announce Type: replace Abstract: Arc-standard derivations over projective dependency trees can be interpreted as the incremental construction of lexicalized ordered trees with conti

researcharxiv-cs-cl
28 May 2026
Applications

A Wolf in Sheep's Clothing: Targeted Routing Hijacking in Federated RAG

DGX agent

arXiv:2605.28112v1 Announce Type: cross Abstract: Federated Retrieval-Augmented Generation (FedRAG) is attractive for privacy-sensitive applications because raw data remain local. As a result, routing

applicationsarxiv-cs-cl
28 May 2026
Safety

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection

DGX agent

arXiv:2605.28664v1 Announce Type: cross Abstract: Safety detection models require examples of HHH (Helpful, Harmless, Honest)-violating outputs for robust generalization, however such examples are sca

safetyarxiv-cs-cl
28 May 2026
Model Releases

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates

DGX agent

arXiv:2605.28440v1 Announce Type: new Abstract: DPO has become a widely adopted alternative to RLHF for aligning LLMs with human preferences, eliminating the need for a separate reward model or RL loo

model-releasesarxiv-cs-cl
28 May 2026
← Previous
1…6869707172…161
Next →