AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
29 May 2026

Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance

SafetyDGX agent

arXiv:2411.14279v2 Announce Type: replace-cross Abstract: Large vision-language models (LVLMs) have achieved impressive results in various vision-language tasks. However, despite showing promising per

MAGA-Bench: Machine-Augment-Generated Text via Alignment Detection Benchmark

Model ReleasesDGX agent

arXiv:2601.04633v2 Announce Type: replace Abstract: Machine-Generated Text (MGT) is becoming increasingly difficult to distinguish from Human-Written Text (HWT). This trend has exacerbated malicious a

Mask the Target: A Plug-and-Play Regularizer Against LoRA Forgetting

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.29498v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become one of the most widely used fine-tuning mechanisms for adapting large language models to new domains, tasks, and u

MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models

Model ReleasesDGX agent

arXiv:2605.28825v1 Announce Type: new Abstract: Large language models (LLMs) frequently encode factual and reasoning knowledge in their internal representations that is not faithfully reflected in the

MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables

ResearchDGX agent

arXiv:2605.29859v1 Announce Type: cross Abstract: Recent speech language models rely on encoders that are optimized separately from autoregressive models. Since these encoders are unaware of the downs

Metric-Dependent Annotation Saturation for Learning from Label Distributions

SafetyDGX agent

arXiv:2605.29797v1 Announce Type: new Abstract: When annotators disagree on a label, the disagreement itself carries signal -- and the number of annotators needed to capture it depends on the evaluati

MIC: Maximizing Informational Capacity in Adaptive Representations via Isotropic Subspace Alignment

SafetyDGX agent

arXiv:2605.29987v1 Announce Type: cross Abstract: Although multi-scales representation learning enables elastic-dimension embeddings, nested subspaces often suffer from dimensional redundancy and spec

Mindscape-Aware Retrieval Augmented Generation for Improved Long Context Understanding

ResearchDGX agent

arXiv:2512.17220v2 Announce Type: replace Abstract: Humans understand long and complex texts by relying on a holistic semantic representation of the content. This global view helps organize prior know

Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs

Model ReleasesDGX agent

arXiv:2605.29737v1 Announce Type: cross Abstract: LLM-based coding assistants are seeing rapid adoption, offering substantial gains in developer productivity. As organizations increasingly ship code t

Mining or Synthesis? Rethinking Exploration Efficiency in Iterative Alignment of Mathematical Reasoning

SafetyDGX agent

arXiv:2602.05370v3 Announce Type: replace Abstract: Iterative Direct Preference Optimization (DPO) has emerged as a widely used paradigm for aligning Large Language Models on reasoning tasks. Existing

Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context

ResearchDGX agent

arXiv:2510.06182v2 Announce Type: replace Abstract: A key component of in-context reasoning is the ability of language models (LMs) to bind entities for later retrieval. For example, an LM might repre

Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels

ResearchDGX agent

arXiv:2605.29800v1 Announce Type: new Abstract: LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a frame

Offloading Score: Measuring AI Reliance Through Counterfactual Workflows

AgentsDGX agent

arXiv:2605.29392v1 Announce Type: cross Abstract: AI tools are increasingly integrated into real-world workflows. However, existing measures of reliance on these tools focus on AI output adoption or o

On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training

ResearchDGX agent

arXiv:2605.29496v1 Announce Type: new Abstract: Post-training has greatly improved reasoning in frontier vision-language models, yet its gains for perception remain comparatively limited, creating a b

PatchBoard: Schema-Grounded State Mutation for Reliable and Auditable LLM Multi-Agent Collaboration

AgentsDGX agent

arXiv:2605.29313v1 Announce Type: new Abstract: LLM multi-agent systems often coordinate through natural-language dialogue or loosely structured shared memory, making intermediate state difficult to v

PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

SafetyDGX agent

arXiv:2605.29582v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown promise as educational tutors, yet effective tutoring requires more than solving problems: it must provide pro

Procedural Pretraining: Warming Up Language Models with Abstract Data

ResearchDGX agent

arXiv:2601.21725v2 Announce Type: replace Abstract: Pretraining language models directly on web-scale corpora is the de facto paradigm. We study an alternative where the model is initially exposed to

Prompt-Level Reward Specifications for Open-Ended Post-Training

Local AiDGX agent

arXiv:2605.29275v1 Announce Type: new Abstract: Open-ended post-training benefits from rewards that make prompt-specific success conditions explicit, rather than relying only on post-hoc scalar scores

Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization

Model ReleasesDGX agent

arXiv:2605.29327v1 Announce Type: new Abstract: Efficient Distillation (EDistill) compresses large language models (LLMs) by structured pruning parameters and tuning lightweight modules with high trai

Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models

TutorialsDGX agent

arXiv:2605.28913v1 Announce Type: new Abstract: Large reasoning models (LRMs) often generate extensive chain-of-thought (CoT) traces before producing a final answer. As explicit textual artifacts, the

Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs

Model ReleasesDGX agent

arXiv:2605.30021v1 Announce Type: new Abstract: Many open-ended instructions have multiple valid answers that users can benefit from seeing, but post-training often narrows an LLM's output space towar

Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents

SafetyDGX agent

arXiv:2605.29447v1 Announce Type: cross Abstract: While GUI agents have advanced rapidly, they often lack the robustness to recover from their own errors, hindering real-world deployment. To bridge th

Resolution Diagnostics for Paired LLM Evaluation

ResearchDGX agent

arXiv:2605.30315v1 Announce Type: new Abstract: Across two public LLM leaderboards, many displayed pairwise rankings do not meet a conventional paired-test resolution target under the actual paired ev

Rethinking Stepwise Model Routing: A Cost-Efficient Table Reasoning Perspective

ResearchDGX agent

arXiv:2605.29319v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on table reasoning tasks but incur substantial inference cost due to long reasoning traces. Ste

Revisiting Observation Reduction for Web Agents: Comprehensive Evaluation with a Lightweight Framework

AgentsDGX agent

arXiv:2605.29397v1 Announce Type: new Abstract: HTML observations in LLM-based web agents are extremely long, and while many reduction methods have been proposed, it remains unclear which methods redu

RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment

Model ReleasesDGX agent

arXiv:2605.28827v1 Announce Type: new Abstract: Open Arabic large language models split into two classes: sub-1B multilingual models that treat Arabic as an afterthought (Qwen2.5-0.5B, Falcon-H1-0.5B)

RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains

SafetyDGX agent

arXiv:2605.29156v1 Announce Type: cross Abstract: Pointwise reward modeling offers critical signals for LLM post-training, yet struggles with absolute scoring in subjective, non-verifiable settings. R

Scaling Laws for Agent Harnesses via Effective Feedback Compute

Model ReleasesDGX agent

arXiv:2605.29682v1 Announce Type: new Abstract: Agent harnesses increasingly determine the performance of language-model systems by deciding how models call tools, receive feedback, verify intermediat

SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?

AgentsDGX agent

arXiv:2605.30104v1 Announce Type: new Abstract: Widely used language-model benchmarks are increasingly saturated, with frontier systems often receiving near-tied scores that standard metrics cannot re

ShapleyLaw: A Game-Theoretic Approach to Multilingual Scaling Laws

ResearchDGX agent

arXiv:2603.17945v2 Announce Type: replace Abstract: In multilingual pretraining, the test loss of a pretrained model is heavily influenced by the proportion of each language in the pretraining data, n

Slogans or Stance? A Label-Light Diagnostic for Entrepreneurial-Discourse Measurement on Chinese SOE Speeches

ResearchDGX agent

arXiv:2605.29188v1 Announce Type: new Abstract: Dictionary methods, topic models, and embedding-similarity scorers are widely used in CSS and management research to measure constructs such as 'entrepr

Spurious Prompts: Can Irrelevant Prompts Steer Large Language Models?

ResearchDGX agent

arXiv:2605.29678v1 Announce Type: new Abstract: Large language models are highly sensitive to prompts, but this sensitivity is usually studied through task-relevant instructions, demonstrations, or re

STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments

Model ReleasesDGX agent

arXiv:2605.29324v1 Announce Type: new Abstract: Mobile GUI agents excel at immediate reactive control but frequently fail in realistic, long-horizon tasks that require memory. This failure stems from

Text-Preserving Lossy Text Compression: A Study of Strategic Deletion and LLM Reconstruction

Model ReleasesDGX agent

arXiv:2605.29000v1 Announce Type: new Abstract: Traditional lossless text compression preserves every byte, but its gains on natural language are often modest in realistic operating regimes. We study

The Anatomy of Conversational Scams: A Topic-Based Red Teaming Analysis of Multi-Turn Interactions in LLMs

SafetyDGX agent

arXiv:2601.03134v2 Announce Type: replace Abstract: As LLMs gain persuasive capabilities through extended dialogues, they create new opportunities for studying adversarial conversational behavior in e

The Trust Paradox: How CS Researchers Engage LLM Leaderboards

Model ReleasesDGX agent

arXiv:2605.28966v1 Announce Type: new Abstract: Large language model (LLM) leaderboards rank AI models using standardized benchmarks and have become highly visible across computer science, despite kno

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2602.15382v2 Announce Type: replace Abstract: Multi-Agent Systems (MAS) powered by Large Language Models have unlocked advanced collaborative reasoning, yet they remain bottlenecked by discrete

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge

ResearchDGX agent

arXiv:2505.16178v2 Announce Type: replace Abstract: While fine-tuning is the standard for injecting factual knowledge into large language models (LLMs), the mechanisms enabling reliable fact recall vi

Understanding Safety-Sensitive Expert Behavior in Mixture-of-Experts LLMs

Model ReleasesDGX agent

arXiv:2605.29708v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) LLMs rely on sparse, router-driven expert activation, yet how safety alignment interacts with routed expert specialization rema

Understanding the Ability of LLMs to Handle Character-Level Perturbation

ResearchDGX agent

arXiv:2510.14365v4 Announce Type: replace Abstract: This work investigates the resilience of contemporary large language models (LLMs) against frequent character-level perturbations. We examine three

UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering

ResearchDGX agent

arXiv:2605.30076v1 Announce Type: new Abstract: Activation-based control steers large language models (LLMs) by intervening on their internal representations during inference, and has emerged as an ef

User-Aware Active Knowledge Acquisition for Emotional Support Dialogue

SafetyDGX agent

arXiv:2605.29715v1 Announce Type: new Abstract: Emotional support plays an important role in dialogue systems, and its success depends on adapting to a user's evolving and implicit needs across multi-

Valency Classification of Mapudungun Verbal Roots. Established by the language's own morphotactics

ResearchDGX agent

arXiv:2604.00789v3 Announce Type: replace Abstract: In the previous work, a lexical (re)categorisation -- or confirmation of the given category -- of roots identified as verbal was undertaken to deter

ValueFlow: Measuring the Propagation of Value Perturbations in Multi-Agent LLM Systems

SafetyDGX agent

arXiv:2602.08567v2 Announce Type: replace-cross Abstract: Multi-agent large language model (LLM) systems increasingly consist of agents that observe and respond to one another's outputs. While value a

Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering

Model ReleasesDGX agent

arXiv:2605.29648v1 Announce Type: new Abstract: Applying reinforcement learning to improve factual accuracy in knowledge-intensive question answering faces a reward design dilemma. Response-level rewa

VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents

Model ReleasesDGX agent

arXiv:2605.30256v1 Announce Type: cross Abstract: Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonve

WaterSearch: A Quality-Aware Search-based Watermarking Framework for Large Language Models

ResearchDGX agent

arXiv:2512.00837v2 Announce Type: replace Abstract: Watermarking acts as a critical safeguard in text generated by Large Language Models (LLMs). By embedding identifiable signals into model outputs, w

What are They Thinking? Delineation, Probing and Tracking of Concepts in LLMs

ResearchDGX agent

arXiv:2605.28823v1 Announce Type: new Abstract: As the influence of LLMs expands, it is imperative to gain insight into their decisions. One way to do that is to develop probes that detect the presenc

What Exactly do Children Receive in Language Acquisition? A Case Study on CHILDES with Automated Detection of Filler-Gap Dependencies

TutorialsDGX agent

arXiv:2603.02082v2 Announce Type: replace Abstract: Children's acquisition of filler-gap dependencies has been argued by some to depend on innate grammatical knowledge, while others suggest that the d

When RL Suppresses Its Own Vocabulary: Recovering Reasoning Diversity in Puzzle-to-Math Transfer

ResearchDGX agent

arXiv:2605.29190v1 Announce Type: cross Abstract: Reinforcement learning using verifiable rewards (RLVR) improves LLM reasoning, but the conditions under which it transfers across domains -- and why i

When the Same Coefficients Reach Different Places: Asymmetric Realizability in Transplanting Tokenizers across Large Language Models

Model ReleasesDGX agent

arXiv:2601.00065v3 Announce Type: replace-cross Abstract: Tokenizer transplant in cross-vocabulary model composition reconstructs donor-only embedding rows as weighted combinations over shared lexical

Who Am I? History-Aware Profiles for Student Simulation in Tutoring Dialogues

ApplicationsDGX agent

arXiv:2605.30051v1 Announce Type: new Abstract: A key part of developing large language model (LLM)-powered, automated tutoring tools is student simulation, i.e., using LLMs to role-play as students,

World Models in Words: Auditing Physical State-Transition Commitments in Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.29585v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used to answer questions about physical scenes, yet most evaluations reduce performance to a final answer

WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction

AgentsDGX agent

arXiv:2605.29341v1 Announce Type: cross Abstract: Multimodal large language models are increasingly deployed as long-horizon agents, where memory must do more than recall: it must track an evolving wo

X-GS: An Extensible Framework for Perceiving and Thinking via 3D Gaussian Splatting

ResearchDGX agent

arXiv:2603.09632v3 Announce Type: replace-cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful technique for novel view synthesis, subsequently extending into numerous spatial AI app

28 May 2026

A new semantically annotated corpus with syntactic-semantic and cross-lingual senses

ResearchDGX agent

arXiv:2605.28494v1 Announce Type: new Abstract: We describe a new sense-tagged corpus for word sense disambiguation. The corpus is constituted of instances of 20 French polysemous verbs. Each verb ins

A tree interpretation of arc standard dependency derivation

ResearchDGX agent

arXiv:2603.27459v2 Announce Type: replace Abstract: Arc-standard derivations over projective dependency trees can be interpreted as the incremental construction of lexicalized ordered trees with conti

A Wolf in Sheep's Clothing: Targeted Routing Hijacking in Federated RAG

ApplicationsDGX agent

arXiv:2605.28112v1 Announce Type: cross Abstract: Federated Retrieval-Augmented Generation (FedRAG) is attractive for privacy-sensitive applications because raw data remain local. As a result, routing

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection

SafetyDGX agent

arXiv:2605.28664v1 Announce Type: cross Abstract: Safety detection models require examples of HHH (Helpful, Harmless, Honest)-violating outputs for robust generalization, however such examples are sca

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates

Model ReleasesDGX agent

arXiv:2605.28440v1 Announce Type: new Abstract: DPO has become a widely adopted alternative to RLHF for aligning LLMs with human preferences, eliminating the need for a separate reward model or RL loo

← Previous
1…5455565758…129
Next →