AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlog
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Agents

Retrieval as Reasoning: Self-Evolving Agent-Native Retrieval via LLM-Wiki

DGX agent

arXiv:2605.25480v1 Announce Type: new Abstract: LLM agents require retrieval to behave less like one-shot context fetching and more like reasoning: searching, reading, traversing, and deciding when ev

agentsarxiv-cs-cl
26 May 2026
Research
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Retrieved In-Context Principles from Previous Mistakes

DGX agent

arXiv:2407.05682v2 Announce Type: replace Abstract: In-context learning (ICL) has been instrumental in adapting Large Language Models (LLMs) to downstream tasks using correct input-output examples. Re

researcharxiv-cs-cl
26 May 2026
Research

ROC Analysis for Evaluating Translation Quality Estimation Systems

DGX agent

arXiv:2605.24721v1 Announce Type: new Abstract: The increasing use of automated translation quality estimation (QE) systems calls for practical, decision-oriented methods for evaluating their performa

researcharxiv-cs-cl
26 May 2026
Model Releases

RotMoLE: Enhancing Mixture of Low-Rank Experts through Rotational Gating Mechanism

DGX agent

arXiv:2605.25565v1 Announce Type: cross Abstract: While Large Language Models (LLMs) are commonly fine-tuned to handle domain-specific tasks before being applied to vertical applications, adapting the

model-releasesarxiv-cs-cl
26 May 2026
Safety

RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry

DGX agent

arXiv:2605.24817v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures have become an increasingly important paradigm for scaling Large Language Models (LLMs). As MoE models are incr

safetyarxiv-cs-cl
26 May 2026
Research

Rubato: Transcribing Piano Music with Timestamps

DGX agent

arXiv:2605.24291v1 Announce Type: cross Abstract: We consider the conversion of musical recordings into human-readable sheet music annotated with timestamps. Such output lets a listener clearly visual

researcharxiv-cs-cl
26 May 2026
Research

Scaling Natural-Language Graph-Based Test Time Compute for Automated Theorem Proving

DGX agent

arXiv:2503.11657v3 Announce Type: replace Abstract: Large language models have demonstrated remarkable capabilities in natural language processing tasks requiring multi-step logical reasoning capabili

researcharxiv-cs-cl
26 May 2026
Model Releases

Schema-Grounded LLM Extraction for FHIR Patient Digital Twins

DGX agent

arXiv:2601.05847v2 Announce Type: replace Abstract: We revisit the problem of constructing interoperable patient digital twins from unstructured electronic health records (EHRs) and argue that the tas

model-releasesarxiv-cs-cl
26 May 2026
Safety

SEAL: Synergistic Co-Evolution of Agents and Learning Environments

DGX agent

arXiv:2605.24426v1 Announce Type: new Abstract: Large Language Model (LLM) agents are increasingly improved through interaction, yet most self-evolution methods adapt either the policy or the learning

safetyarxiv-cs-cl
26 May 2026
Safety

Selective Latent Thinking: Adaptive Compression of LLM Reasoning Chains

DGX agent

arXiv:2605.25745v1 Announce Type: new Abstract: Explicit chain-of-thought (CoT) reasoning substantially improves the reasoning ability of large language models (LLMs), but incurs high inference cost d

safetyarxiv-cs-cl
26 May 2026
Model Releases

SLAP: Stratified Loss-based Pruning for On-Policy Data-Efficient Instruction Tuning

DGX agent

arXiv:2605.23969v1 Announce Type: new Abstract: Instruction tuning has optimized the specialized capabilities of large language models (LLMs), but it often requires extensive datasets and prolonged tr

model-releasesarxiv-cs-cl
26 May 2026
Safety

SliceWorld: A Predictive and Controllable World-State Model for CT Report Generation

DGX agent

arXiv:2605.24371v1 Announce Type: cross Abstract: CT report generation (CTRG) requires models to summarize three-dimensional anatomical context and pathological findings from hundreds of axial slices.

safetyarxiv-cs-cl
26 May 2026
Research

Spiking the training data to correct for test set contamination

DGX agent

arXiv:2605.24818v1 Announce Type: cross Abstract: The literature on test set contamination largely focuses on detection, but the correction of contaminated test scores is underexplored. Our core propo

researcharxiv-cs-cl
26 May 2026
Research

StepGap: A Hybrid NLI-LLM Checker for Step-Level Evidence-Gap Detectionin Multi-Hop Question Answering

DGX agent

arXiv:2605.24733v1 Announce Type: new Abstract: We present extbf{StepGap}, a hybrid NLI-LLM decision tree that detects step-level evidence gaps in multi-hop QA and emits one of three typed labels: ext

researcharxiv-cs-cl
26 May 2026
Agents

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models

DGX agent

arXiv:2605.26014v1 Announce Type: cross Abstract: Many video reasoning tasks require tracking motion, temporal order, and evolving visual states across frames. Existing methods built on large vision-l

agentsarxiv-cs-cl
26 May 2026
Model Releases

StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios

DGX agent

arXiv:2605.25758v1 Announce Type: new Abstract: Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the re

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Structural Abstraction as an Inductive Bias for Non-Stationary Language Model Training

DGX agent

arXiv:2603.17198v2 Announce Type: replace-cross Abstract: A foundational principle in cognitive science holds that intelligent agents do not learn by storing experiences as isolated instances, but by

model-releasesarxiv-cs-cl
26 May 2026
Applications

Structure-Aware RAG: Structured Retrieval Augmented Generation from Noisy Data for Conversational Agents

DGX agent

arXiv:2605.24366v1 Announce Type: new Abstract: Large Language Models (LLMs) have been widely adopted in conversational applications. However, their reliance on parametric knowledge limits reliability

applicationsarxiv-cs-cl
26 May 2026
Model Releases

SURGE: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors

DGX agent

arXiv:2502.11167v5 Announce Type: replace-cross Abstract: Neural surrogate models are powerful and efficient tools in data mining. Meanwhile, large language models (LLMs) have demonstrated remarkable

model-releasesarxiv-cs-cl
26 May 2026
Research

Testing the Deliteralization Hypothesis in Human and Machine Translation

DGX agent

arXiv:2605.25686v1 Announce Type: new Abstract: The recent shift from dedicated NMT systems to general-purpose LLMs has reshaped machine translation, with LLMs reported to produce more fluent, less li

researcharxiv-cs-cl
26 May 2026
Research

Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization

DGX agent

arXiv:2605.25928v1 Announce Type: new Abstract: We describe the winning system for Task 2 of the KSAA-2026 Shared Task on Arabic Speech Dictation with Automatic Diacritization. The task requires produ

researcharxiv-cs-cl
26 May 2026
Model Releases

The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models

DGX agent

arXiv:2605.25510v1 Announce Type: new Abstract: Children increasingly have access to Large Language Models (LLMs), which may expose them to responses that are developmentally inappropriate or require

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

The LSCD Benchmark: a Testbed for Diachronic Word Meaning Tasks

DGX agent

arXiv:2404.00176v3 Announce Type: replace Abstract: Lexical Semantic Change Detection (LSCD) is a complex, lemma-level task, which is usually operationalized based on two subsequently applied usage-le

model-releasesarxiv-cs-cl
26 May 2026
Research

The meaning of prompts and the prompts of meaning: Semiotic reflections and modelling

DGX agent

arXiv:2509.14250v2 Announce Type: replace Abstract: This paper explores prompts and prompting in large language models (LLMs) as dynamic semiotic phenomena, drawing on Peirce's triadic model of signs,

researcharxiv-cs-cl
26 May 2026
Research

The Multilingual Curse at the Retrieval Layer: Evidence from Amharic

DGX agent

arXiv:2605.24556v1 Announce Type: cross Abstract: Multilingual retrieval increasingly underpins cross-lingual question answering and retrieval-augmented generation. Strong zero-shot scores on multilin

researcharxiv-cs-cl
26 May 2026
Research

The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty

DGX agent

arXiv:2605.24718v1 Announce Type: new Abstract: Tokenizer fertility the number of tokens per word imposes a hidden cost on non-English NLP. We measure fertility for ten foundation models across 25 Eur

researcharxiv-cs-cl
26 May 2026
Research

They Are Not the Same: Direct Causes Are Not Grounded Emotion Explanations

DGX agent

arXiv:2605.25208v1 Announce Type: new Abstract: Emotion-Cause Pair Extraction (ECPE) was introduced to explain why an emotion occurs, but this goal is now often reduced to binary pair/non-pair predict

researcharxiv-cs-cl
26 May 2026
Model Releases

TimeSpot: Benchmarking Geo-Temporal Understanding in Vision-Language Models in Real-World Settings

DGX agent

arXiv:2603.06687v2 Announce Type: replace-cross Abstract: Geo-temporal understanding, the ability to infer location, time, and contextual properties from visual input alone, underpins applications suc

model-releasesarxiv-cs-cl
26 May 2026
Agents

Tool-Call Dependency Structure is Linearly Decodable in LLM Agent Residual Streams

DGX agent

arXiv:2605.25310v1 Announce Type: new Abstract: Tool-using LLM agents produce trajectories whose calls form a directed dependency graph: earlier tool outputs supply arguments to later calls. Whether t

agentsarxiv-cs-cl
26 May 2026
Safety

Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content

DGX agent

arXiv:2509.12672v2 Announce Type: replace Abstract: The volume of machine-generated content online has grown dramatically due to the widespread use of Large Language Models (LLMs), leading to new chal

safetyarxiv-cs-cl
26 May 2026
Research

Toxicity in Twitch Chats: An LLM-Based Analysis Across Gaming Communities

DGX agent

arXiv:2605.24000v1 Announce Type: new Abstract: Toxicity in online gaming communities remains a persistent challenge, manifesting across genres, platforms, and player interactions. While much research

researcharxiv-cs-cl
26 May 2026
Model Releases

TRACE: A taxonomy-grounded synthetic dataset for teaching-program generation and session interpretation in Applied Behavior Analysis

DGX agent

arXiv:2605.25038v1 Announce Type: new Abstract: Applied Behavior Analysis (ABA) is a clinical discipline whose documentation, teaching programs and multi-session behavioral logs, is formulaic and high

model-releasesarxiv-cs-cl
26 May 2026
Safety

Trait-Aware Policy Optimization for Autoregressive Multi-Trait Essay Scoring

DGX agent

arXiv:2605.25731v1 Announce Type: new Abstract: Multi-trait essay scoring aims to provide fine-grained evaluation of writing quality across multiple dimensions. However, how to effectively post-train

safetyarxiv-cs-cl
26 May 2026
Applications

Transformers over-extend what humans underlearn: the case of Spanish L-shaped morphome

DGX agent

arXiv:2507.21556v3 Announce Type: replace Abstract: The cognitive reality of irregular morphological patterns has been debated for decades: do speakers extend them to novel forms, or are they lexical

applicationsarxiv-cs-cl
26 May 2026
Applications

Translators as Invisible Teachers of AI: Copyright, Translation Memory, and the Political Economy of Linguistic Data

DGX agent

arXiv:2605.24842v1 Announce Type: new Abstract: This paper examines how the labour of translators has been transformed into foundational data capital for the age of artificial intelligence (AI). Trans

applicationsarxiv-cs-cl
26 May 2026
Research

Triplet-Block Diffusion RWKV

DGX agent

arXiv:2605.25969v1 Announce Type: new Abstract: Causal Transformer language models suffer from strictly sequential decoding and a quadratic per-step attention cost. While linear-time causal models and

researcharxiv-cs-cl
26 May 2026
Model Releases

TypedCSIP: Typed Counterfactual Pretraining for Chinese Legislative Conflict Classification

DGX agent

arXiv:2605.25474v1 Announce Type: new Abstract: TypedCSIP is a typed counterfactual pretraining method for the conflict-classification task of the LCR-CN benchmark (Zhao et al., 2026): given a (superi

model-releasesarxiv-cs-cl
26 May 2026
Research

Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation

DGX agent

arXiv:2605.25903v1 Announce Type: new Abstract: Activation verbalization explains hidden representations in natural language, but existing methods are mostly limited to self-explanation, where each mo

researcharxiv-cs-cl
26 May 2026
Safety

Universal Boosts, Specific Suppressors: Sparse Autoencoder Steering of Medical Vision-Language Models

DGX agent

arXiv:2605.24977v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) often hallucinate findings when generating chest X-ray reports: they fabricate findings that are not present in

safetyarxiv-cs-cl
26 May 2026
Applications

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval

DGX agent

arXiv:2605.24530v1 Announce Type: new Abstract: Document retrieval in real-world scenarios faces significant challenges due to diverse document formats and modalities. Traditional text-based approache

applicationsarxiv-cs-cl
26 May 2026
Local Ai

What Are We Actually Decoding? Source Attribution for Non-Invasive Brain-to-Language Retrieval

DGX agent

arXiv:2605.24524v1 Announce Type: cross Abstract: In non-invasive neural language decoding, results can be inflated by sources that are not stimulus-evoked neural evidence: decoder priors, embedding-b

local-aiarxiv-cs-cl
26 May 2026
Model Releases

What Makes a Medical Checker Trainable? Diagnosing Signal Collapse and Reward Hacking in Checker-Guided RAG for Biomedical QA

DGX agent

arXiv:2605.25988v1 Announce Type: new Abstract: Medical RAG needs evidence-grounded claims, so plugging a claim-level NLI checker into retrieval-augmented RL is intuitive. extbf{We find that the check

model-releasesarxiv-cs-cl
26 May 2026
Research

What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics

DGX agent

arXiv:2510.16435v2 Announce Type: replace-cross Abstract: With the growing use of large language models and conversational interfaces in human-robot interaction, robots' ability to answer user questio

researcharxiv-cs-cl
26 May 2026
Model Releases

When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation

DGX agent

arXiv:2605.25981v1 Announce Type: new Abstract: We document an empirical phenomenon in chain-of-thought and ReAct agents driven by ten large language models from seven architecture families: meaning-b

model-releasesarxiv-cs-cl
26 May 2026
Research

When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift

DGX agent

arXiv:2605.25629v1 Announce Type: new Abstract: Weak-to-strong (W2S) generalization is a promising framework for scalable oversight, yet existing evaluations often test students under matched train--t

researcharxiv-cs-cl
26 May 2026
Safety

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards

DGX agent

arXiv:2605.25864v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewar

safetyarxiv-cs-cl
26 May 2026
Model Releases

WhenLoss: Diagnosing Write and Retrieval Bottlenecks in Long-Context Memory Systems

DGX agent

arXiv:2605.24579v1 Announce Type: new Abstract: Long-context memory systems often fail under fixed budgets, but end-to-end evaluation does not reveal whether evidence was discarded during compression

model-releasesarxiv-cs-cl
26 May 2026
Applications

WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Pretrained Speech Recognition Transformers

DGX agent

arXiv:2509.10452v2 Announce Type: replace Abstract: Pretrained automatic speech recognition (ASR) models such as Whisper perform well but still need domain adaptation to handle unseen parlance. In man

applicationsarxiv-cs-cl
26 May 2026
← Previous
1…7879808182…162
Next →