AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Model Releases

Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure

DGX agent

arXiv:2601.10566v5 Announce Type: replace Abstract: Entity-level unlearning is usually evaluated by what a model says: whether it stops naming the target, refuses a query, or shifts a Truth Ratio dist

model-releasesarxiv-cs-cl
27 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Rethinking the Multilingual Reasoning Gap with Layer Swap

DGX agent

arXiv:2605.26735v1 Announce Type: new Abstract: Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, even when prompted in non-English languages. Prior wor

model-releasesarxiv-cs-cl
27 May 2026
Safety

RICE-PO: Turning Retrieval Interactions into Credit Signals for Reasoning Agents

DGX agent

arXiv:2605.26352v1 Announce Type: new Abstract: Retrieval is increasingly moving from one-shot matching toward interactive reasoning, where language agents iteratively inspect evidence, reformulate qu

safetyarxiv-cs-cl
27 May 2026
Model Releases

Self-Ensembling Vision-Language Models for Chart Data Extraction

DGX agent

arXiv:2605.27298v1 Announce Type: new Abstract: Charts effectively convey quantitative information, but the underlying data are often locked in image form, hindering reuse and analysis. Manually digit

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Self-Verified Distillation: Your Language Model Is Secretly Its Own Synthetic Data Pipeline

DGX agent

arXiv:2605.26132v1 Announce Type: new Abstract: Can post-trained large language models (LLMs) further improve themselves using only unlabeled prompts, without external teachers or feedback from tools?

model-releasesarxiv-cs-cl
27 May 2026
Applications

Semantic Gradients Interactions in SSD: A Case Study in Racial Identity and Hate Speech

DGX agent

arXiv:2605.27322v1 Announce Type: new Abstract: We introduce interaction SSD, an extension of Supervised Semantic Differential that models how semantic meaning varies across moderators such as groups,

applicationsarxiv-cs-cl
27 May 2026
Model Releases

Separating Semantic Competition from Context Length in RAG Reading

DGX agent

arXiv:2605.27294v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems can respond incorrectly even when the correct passage was retrieved. The model must still read the retrieve

model-releasesarxiv-cs-cl
27 May 2026
Tutorials

Share More, Search Less: Collaborative Parallel Thinking for Efficient Test-Time Scaling

DGX agent

arXiv:2605.27030v1 Announce Type: new Abstract: Test-Time Scaling (TTS) enhances the reasoning capabilities of large language models by allocating additional inference compute to explore the solution

tutorialsarxiv-cs-cl
27 May 2026
Model Releases

Shopping Companion: A Memory-Augmented LLM Agent for Real-World E-Commerce Tasks

DGX agent

arXiv:2603.14864v2 Announce Type: replace Abstract: In e-commerce, LLM agents show promise for shopping tasks such as recommendations, budget management, and bundle deals, where accurately capturing u

model-releasesarxiv-cs-cl
27 May 2026
Research

Slide Deck Q&A Quality Assurance App: A Multi-Stage Pipeline for Pedagogical Question Generation

DGX agent

arXiv:2605.26428v1 Announce Type: new Abstract: Generating high-quality, pedagogically useful questions from lecture slide decks is difficult because important instructional content is distributed acr

researcharxiv-cs-cl
27 May 2026
Model Releases

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning

DGX agent

arXiv:2603.28730v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have shown impressive capabilities across diverse tasks, motivating efforts to leverage these models to supervis

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens

DGX agent

arXiv:2508.05305v2 Announce Type: replace Abstract: The recently proposed Large Concept Model (LCM) generates text by predicting a sequence of sentence-level embeddings and training with either mean-s

model-releasesarxiv-cs-cl
27 May 2026
Agents

SPEAR: Code-Augmented Agentic Prompt Optimization

DGX agent

arXiv:2605.26275v1 Announce Type: new Abstract: Automatic prompt engineering (APE) rewrites prompts to improve downstream task performance, but existing APE loops treat the optimizer itself as a fixed

agentsarxiv-cs-cl
27 May 2026
Research

Stylistic Evolution and LLM Neutrality in Singlish Language

DGX agent

arXiv:2601.06580v2 Announce Type: replace Abstract: Singlish is a creole rooted in Singapore's multilingual environment that continues to evolve alongside social and technological change. We examine d

researcharxiv-cs-cl
27 May 2026
Model Releases

SWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue Resolution

DGX agent

arXiv:2603.01327v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strong performance on self-contained programming tasks. However, they still struggle with repository-leve

model-releasesarxiv-cs-cl
27 May 2026
Agents

Telenor Nordics Customer Service self-help corpus

DGX agent

arXiv:2605.26891v1 Announce Type: new Abstract: This paper presents a multilingual customer service self-help corpus comprising 1,122 manually validated documents in Finnish, Danish, Norwegian, and Sw

agentsarxiv-cs-cl
27 May 2026
Model Releases

Temporal Simultaneity Predicts Annotation Quality in Sentiment Corpora

DGX agent

arXiv:2605.27239v1 Announce Type: new Abstract: Annotation quality is difficult to sustain when campaigns span weeks or months with small annotator pools. We present a Setswana sentiment dataset of 3,

model-releasesarxiv-cs-cl
27 May 2026
Safety

The Coverage Illusion: From Pre-retrieval Routing Failure to Post-retrieval Cascades in a Production RAG System

DGX agent

arXiv:2605.27220v1 Announce Type: new Abstract: In modern RAG pipelines, query augmentation methods such as HyDE and query expansion are applied to every query, resulting in substantial LLM inference

safetyarxiv-cs-cl
27 May 2026
Research

The Daily Dose: Workflow-Integrated Large Language Model Automation for Clinical Summarization and Trial Identification in Radiation Oncology

DGX agent

arXiv:2605.26346v1 Announce Type: new Abstract: Objective: To describe the design and early clinical evaluation of The Daily Dose (TDD), an LLM-driven, automated clinical summarization and clinical-tr

researcharxiv-cs-cl
27 May 2026
Agents

The Need for an External Observer Formalizing the Sufficiency Gap: A Mathematical Extension of Mixture Identifiability and Contextual Grounding in Sequence Models

DGX agent

arXiv:2605.26711v1 Announce Type: new Abstract: We construct a binary mixed-regime process with one deterministic textual regime and one random regime governed by an unobserved latent state. Even an i

agentsarxiv-cs-cl
27 May 2026
Safety

To model human linguistic prediction, make LLMs less superhuman

DGX agent

arXiv:2510.05141v2 Announce Type: replace Abstract: When we read, we make predictions about upcoming words; these predictions influence our reading behavior. The success of large language models (LLMs

safetyarxiv-cs-cl
27 May 2026
Tutorials

Towards Just-in-Time Adaptive Feedback: Enhancing Student Learning via Knowledge-Grounded LLM

DGX agent

arXiv:2605.26405v1 Announce Type: new Abstract: Educational interventions are effective tools for enhancing student learning. While Large Language Models (LLMs) allow for generating adaptive feedback

tutorialsarxiv-cs-cl
27 May 2026
Applications

UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action

DGX agent

arXiv:2510.17790v3 Announce Type: replace-cross Abstract: Computer-use agents face a fundamental limitation. They rely exclusively on primitive GUI actions (click, type, scroll), creating brittle exec

applicationsarxiv-cs-cl
27 May 2026
Research

Uncertainty-Aware Budget Allocation for Adaptive Test-Time Reasoning

DGX agent

arXiv:2605.26849v1 Announce Type: new Abstract: Sampling multiple responses improves language model reasoning, but uniform compute allocation is inefficient: easy questions are over-sampled while hard

researcharxiv-cs-cl
27 May 2026
Model Releases

Vectors Are Not Neutral: Sensitive-Information Inference from Exported LLM Representations in Summarization

DGX agent

arXiv:2605.26433v1 Announce Type: new Abstract: Large language model (LLM) summarization systems may pass compact vector representations of private inputs to downstream retrieval, monitoring, audit, o

model-releasesarxiv-cs-cl
27 May 2026
Safety

VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models

DGX agent

arXiv:2510.17759v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) extend large language models with visual reasoning, but their multimodal design also introduces new, underexplor

safetyarxiv-cs-cl
27 May 2026
Research

Verilog-Evolve: Feedback-Driven and Skill-Evolving Verilog Generation

DGX agent

arXiv:2605.26498v1 Announce Type: new Abstract: Large language models (LLMs) have improved Verilog generation from natural-language specifications, but most pipelines still treat generation as isolate

researcharxiv-cs-cl
27 May 2026
Research

When Does Demographic Information Help? Data and Modeling Regimes for Perspective-Aware Hate Speech Detection

DGX agent

arXiv:2605.27313v1 Announce Type: new Abstract: Demographic information is often used to model annotator perspectives in subjective tasks such as hate speech detection, but its benefit is inconsistent

researcharxiv-cs-cl
27 May 2026
Model Releases

Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis

DGX agent

arXiv:2605.26655v1 Announce Type: new Abstract: Automated prompt optimization methods (e.g., DSpy, TextGrad) can substantially improve the performance of large language model (LLM), however, their gen

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

A Comprehensive Dataset for Human vs. AI Generated Text Detection

DGX agent

arXiv:2510.22874v3 Announce Type: replace Abstract: The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authentic

model-releasesarxiv-cs-cl
26 May 2026
Applications

A Lightweight Hybrid Transformer-CRF Architecture for Multi-Type Bangla Medical Entity Recognition

DGX agent

arXiv:2605.25463v1 Announce Type: new Abstract: MedER refers to the identification of medical entities. It is crucial for extracting structured clinical information from unstructured medical text. Man

applicationsarxiv-cs-cl
26 May 2026
Model Releases

A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks

DGX agent

arXiv:2605.23977v1 Announce Type: new Abstract: This paper audits benchmark evaluation in clinical-interview depression detection through four complementary probes across DAIC/E-DAIC, CMDC, ANDROIDS,

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays

DGX agent

arXiv:2605.25652v1 Announce Type: new Abstract: Free-form legal essay evaluation in NLP treats expert inter-rater stability as a single ceiling number, and treats LLM-judge agreement with that ceiling

model-releasesarxiv-cs-cl
26 May 2026
Agents

Act or Clarify? Modeling Sensitivity to Uncertainty and Cost in Communication

DGX agent

arXiv:2602.02843v3 Announce Type: replace Abstract: When deciding how to act under uncertainty, agents may choose to act to reduce uncertainty or they may act despite that uncertainty. In communicativ

agentsarxiv-cs-cl
26 May 2026
Safety

Adaptive Preference Optimization with Uncertainty-aware Utility Anchor

DGX agent

arXiv:2509.10515v1 Announce Type: cross Abstract: Offline preference optimization methods are efficient for large language models (LLMs) alignment. Direct Preference optimization (DPO)-like learning,

safetyarxiv-cs-cl
26 May 2026
Model Releases

AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue

DGX agent

arXiv:2605.23974v1 Announce Type: new Abstract: Current language models create two safety challenges: risk must be detected early enough to avoid exposing harmful continuation, and the harmfulness its

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

DGX agent

arXiv:2505.24876v2 Announce Type: replace-cross Abstract: Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understandi

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios

DGX agent

arXiv:2508.19988v3 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved high accuracy on complex commonsense and mathematical problems that involve the composition of multiple r

model-releasesarxiv-cs-cl
26 May 2026
Hardware

AgentIR: A Workload-Adaptive Cascade Retrieval Substrate for Long-Term Conversational Memory

DGX agent

arXiv:2605.25092v1 Announce Type: cross Abstract: Long-term conversational memory is a retrieval workload classical IR was not built for: the index grows during the query stream, query types shift int

hardwarearxiv-cs-cl
26 May 2026
Model Releases

An Effective-Rank Audit of Alignment-Induced Activation Shifts: Confound Control, Constructive Calibration, and Limits

DGX agent

arXiv:2605.24583v1 Announce Type: cross Abstract: We audit alignment-induced shifts in residual-stream activations of three open-weight instruction-tuned LLMs (Llama-3.1-8B-Instruct, Gemma-2-9B-it, Qw

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents

DGX agent

arXiv:2605.25971v1 Announce Type: new Abstract: While AI agents demonstrate remarkable capabilities in reasoning and tool use, they remain fundamentally reactive: they compute responses only after exp

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

AstroMind: A High-Fidelity Benchmark for Spacecraft Behavior Reasoning Based on Large Language Models

DGX agent

arXiv:2605.24573v1 Announce Type: new Abstract: Understanding why a spacecraft maneuvers -- rather than simply that it did -- is an increasingly important problem for space domain awareness as Earth o

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

AuthTrace: Diagnosing Evidence Construction in Thematically Dense Single-Author Corpora

DGX agent

arXiv:2605.25382v1 Announce Type: new Abstract: Evidence construction systems--chunk retrieval, agent memory, knowledge-graph traversal, and thematic indexing--are evaluated on separate benchmarks wit

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Automated Benchmark Auditing for AI Agents and Large Language Models

DGX agent

arXiv:2605.26079v1 Announce Type: new Abstract: Modern AI benchmarks operate at a complexity that outpaces traditional verification methods. Tasks authored by domain experts often contain implicit ass

model-releasesarxiv-cs-cl
26 May 2026
Agents

AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery

DGX agent

arXiv:2604.05550v2 Announce Type: replace Abstract: Artificial intelligence research increasingly depends on prolonged cycles of reproduction, debugging, and iterative refinement to achieve State-Of-T

agentsarxiv-cs-cl
26 May 2026
Model Releases

Axis-Aligned Semantics for ODRL: Resolving Dimensional Ambiguity in Policy Constraints

DGX agent

arXiv:2602.19878v3 Announce Type: replace Abstract: The Open Digital Rights Language (ODRL) represents policy constraints as triples of a left operand, an operator, and a value. Several spatial operan

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Benchmarking and Learning Real-World Customer Service Dialogue

DGX agent

arXiv:2510.22143v3 Announce Type: replace Abstract: Existing benchmarks and training pipelines for industrial intelligent customer service (ICS) remain misaligned with real-world dialogue requirements

model-releasesarxiv-cs-cl
26 May 2026
Research

Better, Faster: Harnessing Self-Improvement in Large Reasoning Models

DGX agent

arXiv:2605.24998v1 Announce Type: new Abstract: Self-improvement training enables the large reasoning models (LRMs) to improve themselves by self-generating reasoning trajectories as training data wit

researcharxiv-cs-cl
26 May 2026
← Previous
1…7576777879…162
Next →