AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems

DGX agent

arXiv:2605.04018v1 Announce Type: new Abstract: Reasoning-intensive retrieval aims to surface evidence that supports downstream reasoning rather than merely matching topical similarity. This capabilit

model-releasesarxiv-cs-cl
6 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Tutorials

Retrieving Floods without Floodlights: Topic Models as Binary Classifiers for Extreme Climate Events in German News

DGX agent

arXiv:2605.03450v1 Announce Type: new Abstract: In studies of media coverage of extreme climate events, NLP methods have become indispensable for identifying relevant texts in large news databases. St

tutorialsarxiv-cs-cl
6 May 2026
Research

Revisiting Graph-Tokenizing Large Language Models: A Systematic Evaluation of Graph Token Understanding

DGX agent

arXiv:2605.03514v1 Announce Type: new Abstract: The remarkable success of large language models (LLMs) has motivated researchers to adapt them as universal predictors for various graph tasks. As a wid

researcharxiv-cs-cl
6 May 2026
Model Releases

Robust Language Identification for Romansh Varieties

DGX agent

arXiv:2603.15969v2 Announce Type: replace Abstract: The Romansh language has several regional varieties, called idioms, which sometimes have limited mutual intelligibility. Despite this linguistic div

model-releasesarxiv-cs-cl
6 May 2026
Research

Rose-SQL: Role-State Evolution Guided Structured Reasoning for Multi-Turn Text-to-SQL

DGX agent

arXiv:2605.03720v1 Announce Type: new Abstract: Recent advances in Large Reasoning Models (LRMs) trained with Long Chain-of-Thought have demonstrated remarkable capabilities in code generation and mat

researcharxiv-cs-cl
6 May 2026
Agents

S^2tory: Story Spine Distillation for Movie Script Summarization

DGX agent

arXiv:2605.03244v1 Announce Type: new Abstract: Movie scripts pose a fundamental challenge for automatic summarization due to their non-linear, cross-cut narrative structure, which makes surface-level

agentsarxiv-cs-cl
6 May 2026
Model Releases

Safety and accuracy follow different scaling laws in clinical large language models

DGX agent

arXiv:2605.04039v1 Announce Type: new Abstract: Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

SAM-NER: Semantic Archetype Mediation for Zero-Shot Named Entity Recognition

DGX agent

arXiv:2605.03706v1 Announce Type: new Abstract: Zero-shot Named Entity Recognition (ZS-NER) remains brittle under domain and schema shifts, where unseen label definitions often misalign with a large l

model-releasesarxiv-cs-cl
6 May 2026
Research

Scoring Edit Impact in Grammatical Error Correction via Embedded Association Graphs

DGX agent

arXiv:2604.06573v2 Announce Type: replace Abstract: A Grammatical Error Correction (GEC) system produces a sequence of edits to correct an erroneous sentence. The quality of these edits is typically e

researcharxiv-cs-cl
6 May 2026
Agents

SEAD: Self-Evolving Agent for Multi-Turn Service Dialogue

DGX agent

arXiv:2602.03548v3 Announce Type: replace Abstract: Large Language Models have demonstrated remarkable capabilities in open-domain dialogues. However, current methods exhibit suboptimal performance in

agentsarxiv-cs-cl
6 May 2026
Local Ai

Segmenting Human-LLM Co-authored Text via Change Point Detection

DGX agent

arXiv:2605.03723v1 Announce Type: new Abstract: The rise of large language models (LLMs) has created an urgent need to distinguish between human-written and LLM-generated text to ensure authenticity a

local-aiarxiv-cs-cl
6 May 2026
Research

Semantically Enriching Investor Micro-blogs for Opinion-Aware Emotion Analysis: A Practical Approach

DGX agent

arXiv:2605.03092v1 Announce Type: new Abstract: While sentiment analysis is the staple of financial NLP, capturing the nuances of 'why' behind that sentiment remains a challenge. There have been attem

researcharxiv-cs-cl
6 May 2026
Research

Sentiment Analysis of Indonesian Spotify Reviews Using Machine Learning and BiLSTM

DGX agent

arXiv:2605.03443v1 Announce Type: new Abstract: This paper benchmarks classical machine learning and deep learning approaches for three-class sentiment classification of Indonesian Spotify reviews. Us

researcharxiv-cs-cl
6 May 2026
Safety

SERE: Structural Example Retrieval for Enhancing LLMs in Event Causality Identification

DGX agent

arXiv:2605.03701v1 Announce Type: new Abstract: Event Causality Identification (ECI) requires models to determine whether a given pair of events in a context exhibits a causal relationship. While Larg

safetyarxiv-cs-cl
6 May 2026
Local Ai

SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification

DGX agent

arXiv:2605.03301v1 Announce Type: new Abstract: De-identification of clinical text remains essential for secondary use of electronic health records (EHRs), yet public benchmarks such as i2b2 2006/2014

local-aiarxiv-cs-cl
6 May 2026
Research

Should We Still Pretrain Encoders with Masked Language Modeling?

DGX agent

arXiv:2507.00994v4 Announce Type: replace Abstract: Learning high-quality text representations is fundamental to a wide range of NLP tasks. While encoder pretraining has traditionally relied on Masked

researcharxiv-cs-cl
6 May 2026
Model Releases

Simulated Students in Tutoring Dialogues: Substance or Illusion?

DGX agent

arXiv:2601.04025v2 Announce Type: replace Abstract: Advances in large language models (LLMs) enable many new innovations in education. However, evaluating the effectiveness of new technology requires

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Sparse Memory Finetuning as a Low-Forgetting Alternative to LoRA and Full Finetuning

DGX agent

arXiv:2605.03229v1 Announce Type: new Abstract: Adapting a pretrained language model to a new task often hurts the general capabilities it already had, a problem known as catastrophic forgetting. Spar

model-releasesarxiv-cs-cl
6 May 2026
Research

Steer Like the LLM: Activation Steering that Mimics Prompting

DGX agent

arXiv:2605.03907v1 Announce Type: new Abstract: Large language models can be steered at inference time through prompting or activation interventions, but activation steering methods often underperform

researcharxiv-cs-cl
6 May 2026
Local Ai

Stochastic Attention: Connectome-Inspired Randomized Routing for Expressive Linear-Time Attention

DGX agent

arXiv:2604.00754v2 Announce Type: replace Abstract: The whole-brain connectome of a fruit fly comprises over 130K neurons connected with a probability of merely 0.02%, yet achieves an average shortest

local-aiarxiv-cs-cl
6 May 2026
Model Releases

SURE-RAG: Sufficiency and Uncertainty-Aware Evidence Verification for Selective Retrieval-Augmented Generation

DGX agent

arXiv:2605.03534v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) grounds answers in retrieved passages, but retrieval is not verification: a passage can be topical and still fail t

model-releasesarxiv-cs-cl
6 May 2026
Research

Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers

DGX agent

arXiv:2605.03780v1 Announce Type: cross Abstract: Transformers are effective at inferring the latent task from context via two inference modes: recognizing a task seen during training, and adapting to

researcharxiv-cs-cl
6 May 2026
Safety

TeamUp: Semantic Project Matching and Team Formation for Learning at Scale

DGX agent

arXiv:2605.03237v1 Announce Type: cross Abstract: Project-based learning improves student engagement and learning outcomes, yet allocating students to appropriately challenging projects while forming

safetyarxiv-cs-cl
6 May 2026
Research

The Counterexample Game: Iterated Conceptual Analysis and Repair in Language Models

DGX agent

arXiv:2605.03936v1 Announce Type: new Abstract: Conceptual analysis -- proposing definitions and refining them through counterexamples -- is central to philosophical methodology. We study whether lang

researcharxiv-cs-cl
6 May 2026
Research

The Enemy from Within: A Study of Political Delegitimization Discourse in Israeli Political Speech

DGX agent

arXiv:2508.15524v3 Announce Type: replace Abstract: We present the first large-scale computational study of political delegitimization discourse (PDD), defined as symbolic attacks on the normative val

researcharxiv-cs-cl
6 May 2026
Hardware

The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm

DGX agent

arXiv:2505.16932v5 Announce Type: replace-cross Abstract: Computing the polar decomposition and the related matrix sign function has been a well-studied problem in numerical analysis for decades. Rece

hardwarearxiv-cs-cl
6 May 2026
Model Releases

The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It

DGX agent

arXiv:2605.03258v1 Announce Type: cross Abstract: Large language models often fail at simple counting tasks, even when the items to count are explicitly present in the prompt. We investigate whether t

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail

DGX agent

arXiv:2605.03073v1 Announce Type: new Abstract: Niche-domain Indic ASR -- digit strings, currency amounts, addresses, brand names, English/Indic codemix -- is under-served by both open-source SOTA and

model-releasesarxiv-cs-cl
6 May 2026
Research

To Write or to Automate Linguistic Prompts, That Is the Question

DGX agent

arXiv:2603.25169v2 Announce Type: replace Abstract: LLM performance is highly sensitive to prompt design, yet whether automatic prompt optimization can replace expert prompt engineering in linguistic

researcharxiv-cs-cl
6 May 2026
Safety

TRACE: A Metrologically-Grounded Engineering Framework for Trustworthy Agentic AI Systems in Operationally Critical Domains

DGX agent

arXiv:2605.03838v1 Announce Type: new Abstract: We introduce TRACE, a cross-domain engineering framework for trustworthy agentic AI in operationally critical domains. TRACE combines a four-layer refer

safetyarxiv-cs-cl
6 May 2026
Local Ai

Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection

DGX agent

arXiv:2605.02958v1 Announce Type: cross Abstract: Representation Engineering typically relies on static refusal vectors derived from terminal representations. We move beyond this paradigm, demonstrati

local-aiarxiv-cs-cl
6 May 2026
Research

Transformers with Selective Access to Early Representations

DGX agent

arXiv:2605.03953v1 Announce Type: cross Abstract: Several recent Transformer architectures expose later layers to representations computed in the earliest layers, motivated by the observation that low

researcharxiv-cs-cl
6 May 2026
Model Releases

TriBench-Ko: Evaluating LLM Risks in Judicial Workflows

DGX agent

arXiv:2605.03792v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into legal workflows. However, existing benchmarks primarily address proxy tasks, such as bar e

model-releasesarxiv-cs-cl
6 May 2026
Research

Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference

DGX agent

arXiv:2605.03379v1 Announce Type: cross Abstract: Repeated sampling is a standard way to spend test-time compute, but its benefit is controlled by the latent distribution of correctness across example

researcharxiv-cs-cl
6 May 2026
Safety

Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks

DGX agent

arXiv:2502.04419v3 Announce Type: replace-cross Abstract: Generating synthetic datasets via large language models (LLMs) has emerged as a promising approach to improve LLM performance. However, LLMs i

safetyarxiv-cs-cl
6 May 2026
Model Releases

Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development

DGX agent

arXiv:2603.04601v2 Announce Type: replace-cross Abstract: Code generation has emerged as one of AI's highest-impact use cases, yet existing benchmarks measure isolated tasks rather than the complete '

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

When Prompts Interact: Assessing Prompt Arithmetic for Deconfounding under Distribution Shift

DGX agent

arXiv:2605.03096v1 Announce Type: cross Abstract: In classification tasks, models may rely on confounding variables to achieve strong in-distribution performance, capturing spurious features that fail

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

When Should a Language Model Trust Itself? Same-Model Self-Verification as a Conditional Confidence Signal

DGX agent

arXiv:2605.02915v1 Announce Type: new Abstract: Same-model self-verification, prompting a model to audit its own predicted answer, is a plausible confidence signal for selective prediction, but its pr

model-releasesarxiv-cs-cl
6 May 2026
Safety

When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning

DGX agent

arXiv:2605.03314v1 Announce Type: new Abstract: In single-stream autoregressive interfaces, the same tokens both update the model state and constitute an irreversible public commitment. This coupling

safetyarxiv-cs-cl
6 May 2026
Model Releases

Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies

DGX agent

arXiv:2605.03596v1 Announce Type: cross Abstract: Workspace learning requires AI agents to identify, reason over, exploit, and update explicit and implicit dependencies among heterogeneous files in a

model-releasesarxiv-cs-cl
6 May 2026
Research

A framework for analyzing concept representations in neural models

DGX agent

arXiv:2605.01381v1 Announce Type: new Abstract: Understanding how neural models represent human-interpretable concepts is challenging. Prior work has explored linear concept subspaces from diverse per

researcharxiv-cs-cl
5 May 2026
Agents

A Language for Describing Agentic LLM Contexts

DGX agent

arXiv:2605.01920v1 Announce Type: cross Abstract: Large language models are increasingly used within larger systems ('LLM agents'). These make a sequence of LLM calls, each call providing the LLM with

agentsarxiv-cs-cl
5 May 2026
Safety

A Multi-View Media Profiling Suite: Resources, Evaluation, and Analysis

DGX agent

arXiv:2605.01336v1 Announce Type: new Abstract: News outlets shape public opinion at a scale that makes automated detection of political bias and factuality essential. However, the field still lacks u

safetyarxiv-cs-cl
5 May 2026
Model Releases

A multilingual hallucination benchmark: MultiWikiQHalluA

DGX agent

arXiv:2605.02504v1 Announce Type: new Abstract: Most hallucination evaluations focus on English, leaving it unclear whether findings transfer to lower-resource languages. We investigate faithfulness h

model-releasesarxiv-cs-cl
5 May 2026
Research

A Multimodal Dataset for Visually Grounded Ambiguity in Machine Translation

DGX agent

arXiv:2605.02035v1 Announce Type: new Abstract: Ambiguity resolution is a key challenge in multimodal machine translation (MMT), where models must genuinely leverage visual input to map an ambiguous e

researcharxiv-cs-cl
5 May 2026
Model Releases

A Systematic Benchmark of Machine Transliteration Models for the Tajik-Farsi Language Pair: A Comparative Study from Rule-Based to Transformer Architectures

DGX agent

arXiv:2605.02270v1 Announce Type: new Abstract: This paper presents the first comprehensive comparative analysis of modern machine learning architectures for transliteration between Tajik (Cyrillic sc

model-releasesarxiv-cs-cl
5 May 2026
Research

A Systematic Exploration of Text Decomposition and Budget Distribution in Differentially Private Text Obfuscation

DGX agent

arXiv:2605.01065v1 Announce Type: new Abstract: The goal of differentially private text obfuscation is to obfuscate, or 'perturb', input texts with Differential Privacy (DP) guarantees, such that the

researcharxiv-cs-cl
5 May 2026
Safety

A Theoretical Game of Attacks via Compositional Skills

DGX agent

arXiv:2605.01034v1 Announce Type: new Abstract: As large language models grow increasingly capable, concerns about their safe deployment have intensified. While numerous alignment strategies aim to re

safetyarxiv-cs-cl
5 May 2026
← Previous
1…107108109110111…161
Next →