AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use

DGX agent

arXiv:2603.03205v2 Announce Type: replace Abstract: Agentic language models operate in a fundamentally different safety regime than chat models: they must plan, call tools, and execute long-horizon ac

model-releasesarxiv-cs-cl
4 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

LifeSide: Benchmarking Agents as Lifelong Digital Companions

DGX agent

arXiv:2606.04660v1 Announce Type: new Abstract: Lifelong digital companions must integrate cross-session cues, continually update their understanding of users, and adapt to shifting privacy boundaries

model-releasesarxiv-cs-cl
4 Jun 2026
Research

Light or Full Verb? A Minimal-Pair Dataset for Probing Phraseological Competence in Language Models

DGX agent

arXiv:2606.05087v1 Announce Type: new Abstract: Frequent English verbs such as 'have' and 'make' can function either as collocates in light-verb constructions or as full lexical predicates, as in 'mak

researcharxiv-cs-cl
4 Jun 2026
Research

LiSeCo: Linear Semantic Control for Language Generation

DGX agent

arXiv:2405.15454v4 Announce Type: replace Abstract: The prevalence of Large Language Models (LLMs) in critical applications highlights the need for controlled language generation methods that are both

researcharxiv-cs-cl
4 Jun 2026
Safety

Listening to the Workforce: Measuring Construction Worker Safety Attitudes from Social Media Discourse Using LLMs

DGX agent

arXiv:2606.04450v1 Announce Type: new Abstract: Worker safety attitudes are key determinants of whether protective practices are applied or bypassed on construction sites. Yet measuring them at scale

safetyarxiv-cs-cl
4 Jun 2026
Model Releases

LLMs + Persona-Plug = Personalized LLMs

DGX agent

arXiv:2409.11901v2 Announce Type: replace Abstract: Personalization plays a critical role in numerous language tasks and applications, since users with the same requirements may prefer diverse outputs

model-releasesarxiv-cs-cl
4 Jun 2026
Model Releases

Long Live Fine-Tuning: Task-Specific Transformers Outperform Zero-Shot LLMs for Misinformation Response Classification on Reddit

DGX agent

arXiv:2606.04274v1 Announce Type: new Abstract: As large language models (LLMs) become default tools for online information verification, an implicit assumption follows them: that scale and general ca

model-releasesarxiv-cs-cl
4 Jun 2026
Research

Multilingual Long-Form Speech Instruction Following: KIT's Submission to IWSLT 2026

DGX agent

arXiv:2606.04730v1 Announce Type: new Abstract: With the advent of Large Language Models, single-task and token-based multi-task models have evolved into instruction-based systems that infer task and

researcharxiv-cs-cl
4 Jun 2026
Safety

MusaCoder: Native GPU Kernel Generation with Full-Stack Training on Moore Threads GPU

DGX agent

arXiv:2606.04847v1 Announce Type: cross Abstract: Native GPU kernel generation turns high-level tensor programs into executable, efficient low-level code. Existing Large Language Models (LLMs) struggl

safetyarxiv-cs-cl
4 Jun 2026
Model Releases

NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models

DGX agent

arXiv:2606.04773v1 Announce Type: cross Abstract: Reliable evaluation of human motion understanding is fundamental to advancing embodied AI, robotics, and animation. However, existing benchmarks suffe

model-releasesarxiv-cs-cl
4 Jun 2026
Research

Noisy memory encoding explains negative polarity illusions

DGX agent

arXiv:2606.04340v1 Announce Type: new Abstract: A sentence like 'The authors that no critics recommended have ever received acknowledgment for a best-selling novel' is sometimes rated as acceptable ev

researcharxiv-cs-cl
4 Jun 2026
Safety

Off-Distribution Voices: Fanfiction Subgenres as Universal Vernacular Jailbreaks for Aligned LLMs

DGX agent

arXiv:2606.04483v1 Announce Type: new Abstract: Existing jailbreaks against aligned LLMs are discrete artifacts whose surface forms are easy to fingerprint and patch. We argue that the real failure mo

safetyarxiv-cs-cl
4 Jun 2026
Agents

Optimizing the Cost-Quality Tradeoff of Agentic Theorem Provers in Lean

DGX agent

arXiv:2606.04883v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in workflows for generating formal proofs in Lean. These workflows often decompose problems into smal

agentsarxiv-cs-cl
4 Jun 2026
Safety

Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning

DGX agent

arXiv:2601.07408v2 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has emerged as a promising critic-free reinforcement learning paradigm for reasoning tasks. However, stand

safetyarxiv-cs-cl
4 Jun 2026
Model Releases

Parameter-Efficient Fine-Tuning with Learnable Rank

DGX agent

arXiv:2606.04325v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) is a popular parameter-efficient fine-tuning (PEFT) method that restricts weight updates to low-rank adapters, introducing a

model-releasesarxiv-cs-cl
4 Jun 2026
Safety

PersonaTree: Structured Lifecycle Memory for Person Understanding in LLM Agents

DGX agent

arXiv:2606.04780v1 Announce Type: new Abstract: Persistent LLM agents require memory representations that make the formation of person understanding explicit across long term interaction. Existing age

safetyarxiv-cs-cl
4 Jun 2026
Research

Physics-Informed Neural Network Modeling of Biodegradable Contaminant Transport through GCL/SL Composite Liners

DGX agent

arXiv:2606.04392v1 Announce Type: cross Abstract: This study develops a two-domain physics-informed neural network framework for contaminant transport through a GCL/SL composite liner system, in which

researcharxiv-cs-cl
4 Jun 2026
Safety

Probing Outcome-Level Resemblance and Mechanism-Level Alignment in LLM Risk Decisions: Evidence from the St. Petersburg Game

DGX agent

arXiv:2606.04978v1 Announce Type: new Abstract: LLMs can appear cautious in risk decision-making tasks, yet cautious-looking outputs do not necessarily indicate alignment with human decision-making me

safetyarxiv-cs-cl
4 Jun 2026
Model Releases

Prompt-Level Distillation: A Non-Parametric Alternative to Model Fine-Tuning for Efficient Reasoning

DGX agent

arXiv:2602.21103v2 Announce Type: replace Abstract: Advanced reasoning typically requires Chain-of-Thought prompting, which is accurate but incurs prohibitive latency and substantial test-time inferen

model-releasesarxiv-cs-cl
4 Jun 2026
Research

Query-based Cross-Modal Projector Bolstering Mamba Multimodal LLM

DGX agent

arXiv:2606.04719v1 Announce Type: new Abstract: The Transformer's quadratic complexity with input length imposes an unsustainable computational load on large language models (LLMs). In contrast, the S

researcharxiv-cs-cl
4 Jun 2026
Model Releases

RAMPART: Registry-based Agentic Memory with Priority-Aware Runtime Transformation

DGX agent

arXiv:2606.04628v1 Announce Type: new Abstract: RAMPART is a compile-time memory model and pure in-RAM block registry for LLM-based agents. Context assembly is a programmable runtime operation where c

model-releasesarxiv-cs-cl
4 Jun 2026
Research

Read the Trace, Steer the Path: Trajectory-Aware Reinforcement Learning for Diffusion Language Models

DGX agent

arXiv:2606.04396v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) generate responses by iteratively unmasking and revising many positions in parallel. This process leaves a rich

researcharxiv-cs-cl
4 Jun 2026
Research

Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy

DGX agent

arXiv:2606.04680v1 Announce Type: cross Abstract: Automatic speech recognition systems commonly rely on reference transcriptions for evaluation, while reference-free approaches often depend on interna

researcharxiv-cs-cl
4 Jun 2026
Model Releases

Reasoning over Boundaries: Enhancing Specification Alignment via Test-time Deliberation

DGX agent

arXiv:2509.14760v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied in diverse real-world scenarios, each governed by bespoke behavioral and safety specifications

model-releasesarxiv-cs-cl
4 Jun 2026
Safety

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents

DGX agent

arXiv:2606.04703v1 Announce Type: new Abstract: Experience internalization converts contextual experience from past interactions into reusable parametric capability, offering a promising path toward c

safetyarxiv-cs-cl
4 Jun 2026
Research

SAID: Accelerating Diffusion-Based Language Models via Scaffold-Aware Iterative Decoding

DGX agent

arXiv:2606.04974v1 Announce Type: new Abstract: Diffusion large language models (DLLMs) enable non-autoregressive generation by iteratively denoising corrupted token sequences with bidirectional conte

researcharxiv-cs-cl
4 Jun 2026
Research

SANE Schema-aware Natural-language Evaluation of Biological Data

DGX agent

arXiv:2606.04500v1 Announce Type: new Abstract: High-throughput microscopy generates large, structured datasets capturing cellular responses to pharmacological perturbations, but accessing these datas

researcharxiv-cs-cl
4 Jun 2026
Safety

Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing

DGX agent

arXiv:2512.08094v2 Announce Type: replace Abstract: The goal of this work is to develop a universal approach for aligning subtitles (i.e., spoken language text with corresponding timestamps) to contin

safetyarxiv-cs-cl
4 Jun 2026
Local Ai

Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data

DGX agent

arXiv:2606.05122v1 Announce Type: new Abstract: Large language models are increasingly evaluated by other models, raising a natural question: can a model predict how a judge will score its own output?

local-aiarxiv-cs-cl
4 Jun 2026
Local Ai

SemBlock: Semantic Boundary Dynamic Blocks for Diffusion LLMs

DGX agent

arXiv:2606.04964v1 Announce Type: new Abstract: Diffusion language models (DLMs) generate text through iterative denoising, and blockwise decoding improves their practicality by committing tokens in l

local-aiarxiv-cs-cl
4 Jun 2026
Model Releases

SMADE-IE: Sparse Multi-Agent Framework with Evidence-Driven Debate for Zero-Shot Information Extraction

DGX agent

arXiv:2606.04691v1 Announce Type: new Abstract: Zero-shot information extraction (IE) with large language models (LLMs) has attracted increasing attention due to its flexibility in adapting to new sch

model-releasesarxiv-cs-cl
4 Jun 2026
Safety

SocialCoach: Personalized Social Skill Learning with RL-based Agentic Tutoring and Practice

DGX agent

arXiv:2606.04155v1 Announce Type: cross Abstract: Social skills such as negotiation and leadership are crucial for personal and professional success in today's interconnected world. However, scalable

safetyarxiv-cs-cl
4 Jun 2026
Model Releases

Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems

DGX agent

arXiv:2407.03956v3 Announce Type: replace-cross Abstract: Prior research has enhanced the ability of Large Language Models (LLMs) to solve logic puzzles using techniques such as chain-of-thought promp

model-releasesarxiv-cs-cl
4 Jun 2026
Hardware

SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference

DGX agent

arXiv:2606.04511v1 Announce Type: new Abstract: Sparse attention reduces compute and memory bandwidth for long-context LLM inference. However, two key challenges remain: (1) KV cache capacity still gr

hardwarearxiv-cs-cl
4 Jun 2026
Safety

SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space

DGX agent

arXiv:2511.20102v3 Announce Type: replace Abstract: Sparse attention reduces the quadratic complexity of full self-attention but faces two challenges: (1) an attention gap, where applying sparse atten

safetyarxiv-cs-cl
4 Jun 2026
Agents

Stateful Visual Encoders for Vision-Language Models

DGX agent

arXiv:2606.04433v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used in multi-image, multi-turn agentic settings where decisions depend on visual changes. However, in

agentsarxiv-cs-cl
4 Jun 2026
Model Releases

Stepwise Reasoning Enhancement for LLMs via External Subgraph Generation

DGX agent

arXiv:2606.04454v1 Announce Type: new Abstract: Large language models have shown strong performance in natural language generation and downstream reasoning tasks, but they still struggle with logical

model-releasesarxiv-cs-cl
4 Jun 2026
Model Releases

STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations

DGX agent

arXiv:2606.05165v1 Announce Type: cross Abstract: Training Data Attribution (TDA) seeks to trace a model's predictions back to its training data. The gold standard for TDA relies on causal interventio

model-releasesarxiv-cs-cl
4 Jun 2026
Model Releases

TaDA: Calibrated Probe Gating for Task-Domain LoRA Merging

DGX agent

arXiv:2606.05016v1 Announce Type: new Abstract: Combining a task LoRA adapter with a domain LoRA adapter into a single unified model is a practical yet largely unexplored challenge. Existing methods t

model-releasesarxiv-cs-cl
4 Jun 2026
Applications

The Mechanistic Emergence of Symbol Grounding in Language Models

DGX agent

arXiv:2510.13796v3 Announce Type: replace Abstract: Symbol grounding (Harnad, 1990) describes how symbols such as words acquire their meanings by connecting to real-world sensorimotor experiences. Rec

applicationsarxiv-cs-cl
4 Jun 2026
Research

Training-Free Lexical-Dense Fusion for Conversational-Memory Retrieval

DGX agent

arXiv:2606.04194v1 Announce Type: cross Abstract: Retrieving the few past turns that answer a new query across long multi-session histories is the retrieval bottleneck behind long-term conversational

researcharxiv-cs-cl
4 Jun 2026
Research

Translation Heads: Disentangling meaning from language in LLM-based machine translation

DGX agent

arXiv:2602.04613v2 Announce Type: replace Abstract: Mechanistic Interpretability (MI) seeks to explain how neural networks implement their capabilities, but the scale of Large Language Models (LLMs) h

researcharxiv-cs-cl
4 Jun 2026
Research

T^star: Progressive Block Scaling for Masked Diffusion Language Models Through Trajectory Aware Reinforcement Learning

DGX agent

arXiv:2601.11214v5 Announce Type: replace Abstract: We present T^star, a simple TraceRL-based training curriculum for progressive block-size scaling in masked diffusion language models (MDMs). Startin

researcharxiv-cs-cl
4 Jun 2026
Research

UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding

DGX agent

arXiv:2307.00862v3 Announce Type: replace-cross Abstract: Vision-language tasks, such as VQA, SNLI-VE, and VCR are challenging because they require the model's reasoning ability to understand the sema

researcharxiv-cs-cl
4 Jun 2026
Applications

Using Text-Based Causal Inference to Disentangle Factors Influencing Online Review Ratings

DGX agent

arXiv:2606.04286v1 Announce Type: new Abstract: Online reviews provide valuable insights into the perceived quality of facets of a product or service. While aspect-based sentiment analysis has focused

applicationsarxiv-cs-cl
4 Jun 2026
Research

Validity Threats for Foundation Model Research

DGX agent

arXiv:2606.05029v1 Announce Type: cross Abstract: Controlled experiments are the backbone of machine learning research, but at the scale of modern foundation models, they have become prohibitively exp

researcharxiv-cs-cl
4 Jun 2026
Model Releases

VCIFBench: Evaluating Complex Instruction Following for Video Understanding

DGX agent

arXiv:2606.04588v1 Announce Type: new Abstract: Multimodal large language models have made rapid progress in video understanding, yet existing benchmarks largely rely on simple prompts and provide lim

model-releasesarxiv-cs-cl
4 Jun 2026
Safety

VentAgent: When LLMs Learn to Breathe -- Multi-Objective Arbitration for ARDS Ventilation

DGX agent

arXiv:2606.04632v1 Announce Type: cross Abstract: Mechanical ventilation for Acute Respiratory Distress Syndrome (ARDS) requires balancing competing physiological goals, including oxygenation, lung pr

safetyarxiv-cs-cl
4 Jun 2026
← Previous
1…5657585960…161
Next →