AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
5 May 2026

ReMedi: Reasoner for Medical Clinical Prediction

ApplicationsDGX agent

arXiv:2605.01474v1 Announce Type: new Abstract: Predicting future clinical outcomes from electronic health records (EHR) remains challenging due to the complexity and heterogeneity of patient data. LL

Revisiting Semantic Role Labeling: Efficient Structured Inference with Dependency-Informed Analysis

ResearchDGX agent

arXiv:2605.02505v1 Announce Type: new Abstract: Semantic Role Labeling (SRL) provides an explicit representation of predicate-argument structure, capturing linguistically grounded relations such as wh

RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences

Model ReleasesDGX agent

arXiv:2605.01831v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback has become the standard paradigm for language model alignment, where reward models directly determine alignme


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SCARV: Structure-Constrained Aggregation for Stable Sample Ranking in Redundant NLP Datasets

ResearchDGX agent

arXiv:2605.00944v1 Announce Type: cross Abstract: Sample-level rankings are increasingly used in data-centric NLP for analysis, filtering, debugging, and curation, yet existing pipelines typically sco

SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning

Model ReleasesDGX agent

arXiv:2605.01489v1 Announce Type: cross Abstract: Frontier scientific reasoning is rapidly emerging as a key foundation for advancing AI agents in automated scientific discovery. Deep research agents

SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures

Model ReleasesDGX agent

arXiv:2605.02601v1 Announce Type: new Abstract: We present our shared task on evaluating the adaptability of LLMs and NLP systems across multiple languages and cultures. The task data consist of an ex

Sentiment Analysis of Mobile Legends App Reviews Using Machine Learning and LSTM-Based Deep Learning Models

ResearchDGX agent

arXiv:2605.01317v1 Announce Type: new Abstract: This paper compares Machine Learning and LSTM-based Deep Learning methods for sentiment analysis of Mobile Legends app reviews. Using a dataset of 10,00

Shadow-Loom: Causal Reasoning over Graphical World Model of Narratives

Model ReleasesDGX agent

arXiv:2605.02475v1 Announce Type: cross Abstract: Stories hold a reader's attention because they have causes, secrets, and consequences. Shadow-Loom is an experimental open-source framework that turns

Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting

Model ReleasesDGX agent

arXiv:2605.02105v1 Announce Type: cross Abstract: Pretraining optimizers are tuned to produce the strongest possible base model, on the assumption that a stronger starting point yields a stronger mode

SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 25+ Sign Languages

SafetyDGX agent

arXiv:2605.01720v1 Announce Type: cross Abstract: Existing large-scale sign language resources typically provide supervision only at the level of raw video-text alignment and are often produced in lab

Social Dynamics as Critical Vulnerabilities that Undermine Objective Decision-Making in LLM Collectives

AgentsDGX agent

arXiv:2604.06091v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly acting as human delegates in multi-agent environments, where a representative agent integrates di

Spatiotemporal Hidden-State Dynamics as a Signature of Internal Reasoning in Large Language Models

ResearchDGX agent

arXiv:2605.01853v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate extended solutions, yet it remains unclear whether these traces reflect substantive internal computation or merel

SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection

ResearchDGX agent

arXiv:2605.02888v1 Announce Type: cross Abstract: Speculative decoding accelerates large language model (LLM) inference by using a small draft model to propose candidate tokens that a larger target mo

Spoken Language Identification with Pre-trained Models and Margin Loss

Model ReleasesDGX agent

arXiv:2605.01905v1 Announce Type: cross Abstract: For the speaker-controlled spoken language identification task proposed in the TidyLang Challenge 2026, this paper proposes a language identification

SRA: Span Representation Alignment for Large Language Model Distillation

SafetyDGX agent

arXiv:2605.01205v1 Announce Type: new Abstract: Cross-Tokenizer Knowledge Distillation (CTKD) enables knowledge transfer between a large language model and a smaller student, even when they employ dif

SRTJ: Self-Evolving Rule-Driven Training-Free LLM Jailbreaking

Model ReleasesDGX agent

arXiv:2605.00974v1 Announce Type: cross Abstract: LLMs are increasingly equipped with safety alignment mechanisms, yet recent studies demonstrate that they remain vulnerable to jailbreaking attacks th

STAGE: A Full-Screenplay Benchmark for Reasoning over Evolving Storie

Model ReleasesDGX agent

arXiv:2601.08510v3 Announce Type: replace Abstract: Movie screenplays are rich long-form narratives that interleave complex character relationships, temporally ordered events, and dialogue-driven inte

StressEval: Failure-Driven Dynamic Benchmarking for Knowledge-Intensive Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2605.01939v1 Announce Type: new Abstract: Static benchmarks for LLMs are increasingly compromised by contamination and overfitting especially on knowledge intensive reasoning tasks While recent

Structural Dilemmas and Developmental Pathways of Legal Argument Mining in the Era of Artificial Intelligence

ApplicationsDGX agent

arXiv:2605.02308v1 Announce Type: new Abstract: Against the backdrop of rapid advances in artificial intelligence, legal argument mining has emerged as an important research area linking legal texts w

SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation

Model ReleasesDGX agent

arXiv:2508.15658v5 Announce Type: replace Abstract: The rapid growth of academic literature makes the manual creation of scientific surveys increasingly infeasible. While large language models show pr

Synthetic Users, Real Differences: an Evaluation Framework for User Simulation in Multi-Turn Conversations

ResearchDGX agent

arXiv:2605.02624v1 Announce Type: new Abstract: There is growing interest in exploring user simulation as an alternative to gathering and scoring real user-chatbot interactions for AI chatbot evaluati

TagRAG: Tag-guided Hierarchical Knowledge Graph Retrieval-Augmented Generation

Local AiDGX agent

arXiv:2601.05254v3 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation enhances language models by retrieving external knowledge to support informed and grounded responses. However,

TCDA: Thread-Constrained Discourse-Aware Modeling for Conversational Sentiment Quadruple Analysis

Model ReleasesDGX agent

arXiv:2605.01717v1 Announce Type: new Abstract: Conversational Aspect-based Sentiment Quadruple Analysis (DiaASQ) needs to capture the complex interrelationships in multiple rounds of dialogues. Exist

Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines

Model ReleasesDGX agent

arXiv:2605.01077v1 Announce Type: new Abstract: Brazil's Unified Health System (SUS) relies on official clinical guidelines that define diagnostic criteria, treatments, dosages, and monitoring procedu

TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models

Model ReleasesDGX agent

arXiv:2504.20605v2 Announce Type: replace Abstract: Moral stories are a time-tested vehicle for transmitting values, yet modern NLP lacks a large, structured corpus that couples coherent narratives wi

The 2026 ACII Dyadic Conversations (DaiKon) Workshop & Challenge

Model ReleasesDGX agent

arXiv:2605.02672v1 Announce Type: cross Abstract: The 2026 ACII Dyadic Conversations (ACII-DaiKon) Workshop & Challenge introduces a benchmark for modeling interpersonal affect and social dynamics in

The Company You Keep: How LLMs Respond to Dark Triad Traits

ResearchDGX agent

arXiv:2603.04299v3 Announce Type: replace Abstract: Large Language Models (LLMs) often exhibit highly agreeable and reinforcing conversational styles, also known as AI-sycophancy. Although this behavi

The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't

Model ReleasesDGX agent

arXiv:2605.01771v1 Announce Type: new Abstract: An auditor instructs an AI assistant: 'open each file individually using the Read tool -- no scripts, no agents.' The AI replies 'Yes' -- then issues a

The Compliance Trap: How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure

Model ReleasesDGX agent

arXiv:2605.02398v1 Announce Type: cross Abstract: As frontier AI models are deployed in high-stakes decision pipelines, their ability to maintain metacognitive stability -- knowing what they do not kn

The Cylindrical Representation Hypothesis for Language Model Steering

ResearchDGX agent

arXiv:2605.01844v1 Announce Type: new Abstract: Steering is a widely used technique for controlling large language models, yet its effects are often unstable and hard to predict. Existing theoretical

The grip of grammar on meaning uncertainty: cross-linguistic evidence, neural correlates, and clinical relevance

ApplicationsDGX agent

arXiv:2605.01537v1 Announce Type: new Abstract: Isolated word meanings are inherently uncertain. This uncertainty reduces when they are combined and anchored in context. We propose that grammar compre

The Pre-Training Study of Expanded-SPLADE Models on Web Document Titles

ResearchDGX agent

arXiv:2605.01407v1 Announce Type: cross Abstract: Masked Language Modeling (MLM) pre-training is one of the primary ways to initialize Neural Information Retrieval (IR) models prior to retrieval fine-

The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning

AgentsDGX agent

arXiv:2605.01704v1 Announce Type: new Abstract: When copies of the same language model are prompted to debate, they produce diverse phrasings of one perspective rather than diverse perspectives. Multi

Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation

ResearchDGX agent

arXiv:2605.02496v1 Announce Type: cross Abstract: Tibetan text-to-speech (TTS) has long been challenged by scarce speech resources, significant dialectal variation, and the complex mapping between wri

TIME: Temporally Intelligent Meta-reasoning Engine for Context-Triggered Explicit Reasoning

Model ReleasesDGX agent

arXiv:2601.05300v2 Announce Type: replace-cross Abstract: Reasoning-oriented language models typically expose explicit reasoning as a long, front-loaded chain of 'thinking' tokens before the main outp

TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs

Model ReleasesDGX agent

arXiv:2510.15545v4 Announce Type: replace Abstract: Accelerating the inference of large language models (LLMs) has been a critical challenge in generative AI. Speculative decoding (SD) substantially i

Toward Culturally Grounded Natural Language Processing

Model ReleasesDGX agent

arXiv:2603.26013v2 Announce Type: replace Abstract: Multilingual NLP is often treated as a route to global inclusion, but linguistic coverage and cultural competence frequently diverge. This paper syn

VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation

ApplicationsDGX agent

arXiv:2602.21054v2 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) frequently hallucinate, limiting their safe deployment in real-world applications. Existing LLM self-eval

Verbal-R3: Verbal Reranker as the Missing Bridge between Retrieval and Reasoning

AgentsDGX agent

arXiv:2605.01399v1 Announce Type: new Abstract: The conventional Retrieval-Augmented Generation (RAG) paradigm of injecting raw retrieved texts into the Large Language Model (LLM)'s context often resu

VERGE: Formal Refinement and Guidance Engine for Verifiable LLM Reasoning

Local AiDGX agent

arXiv:2601.20055v2 Announce Type: replace Abstract: Despite the syntactic fluency of Large Language Models (LLMs), ensuring their logical correctness in high-stakes domains remains a fundamental chall

VeRO: An Evaluation Harness for Agents to Optimize Agents

Model ReleasesDGX agent

arXiv:2602.22480v2 Announce Type: replace-cross Abstract: An important emerging application of coding agents is agent optimization: the iterative improvement of a target agent through edit-execute-eva

Virtual Speech Therapist: A Clinician-in-the-Loop AI Speech Therapy Agent for Personalized and Supervised Therapy

SafetyDGX agent

arXiv:2605.01101v1 Announce Type: cross Abstract: This paper develops Virtual Speech Therapist (VST), an intelligent agent-based platform that streamlines stuttering assessment and delivers customized

Watermarking LLM Agent Trajectories

Model ReleasesDGX agent

arXiv:2602.18700v2 Announce Type: replace-cross Abstract: LLM agents rely heavily on high-quality trajectory data to guide their problem-solving behaviors, yet producing such data requires substantial

What Single-Prompt Accuracy Misses: A Multi-Variant Reliability Audit of Language Models

Model ReleasesDGX agent

arXiv:2605.02038v1 Announce Type: new Abstract: Single-prompt accuracy is the dominant way to benchmark language models, but it can miss reliability failures that matter. We evaluate a 15-model open-w

When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition

Model ReleasesDGX agent

arXiv:2605.02782v1 Announce Type: cross Abstract: Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech. Recent audio-language models raise the possibility

When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models

Model ReleasesDGX agent

arXiv:2605.02363v1 Announce Type: new Abstract: Deployed language models must produce outputs that are both correct and format-compliant. We study this structured-output reliability gap using two math

When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering

Model ReleasesDGX agent

arXiv:2601.19827v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) extends large language models (LLMs) beyond parametric knowledge, yet it is unclear when iterative retrieval-re

When Less is Enough: Efficient Inference via Collaborative Reasoning

ResearchDGX agent

arXiv:2605.01111v1 Announce Type: cross Abstract: In this work, we introduce DUET (Dual-model Efficient Two-stage inference), a collaborative inference framework in which a capable model and a lightwe

Where Do Prompt Perturbations Break Generation? A Segment-Level View of Robustness in LoRA-Tuned Language Models

ResearchDGX agent

arXiv:2605.01605v1 Announce Type: new Abstract: Large language models are sensitive to minor prompt perturbations, yet existing robustness methods usually enforce consistency at the whole-sequence lev

Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework

SafetyDGX agent

arXiv:2605.01416v1 Announce Type: cross Abstract: The increasing scale and complexity of online platforms raises critical policy questions around harmful content, digital well-being, and user autonomy

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization

HardwareDGX agent

arXiv:2605.02262v1 Announce Type: cross Abstract: Recently, video language models (VLMs) have been applied in various fields. However, the visual token sequence of the VLM is too long, which may cause

Zero-Shot Confidence Estimation for Small LLMs: When Supervised Baselines Aren't Worth Training

Local AiDGX agent

arXiv:2605.02241v1 Announce Type: cross Abstract: How reliably can a small language model estimate its own correctness? The answer determines whether local-to-cloud routing-escalating queries a cheap

4 May 2026

A11y-Compressor: A Framework for Enhancing the Efficiency of GUI Agent Observations through Visual Context Reconstruction and Redundancy Reduction

Model ReleasesDGX agent

arXiv:2605.00551v1 Announce Type: new Abstract: AI agents that interact with graphical user interfaces (GUIs) require effective observation representations for reliable grounding. The accessibility tr

Adaptive Querying with AI Persona Priors

ResearchDGX agent

arXiv:2605.00696v1 Announce Type: cross Abstract: We study adaptive querying for learning user-dependent quantities of interest, such as responses to held-out items and psychometric indicators, within

ADVICE: Answer-Dependent Verbalized Confidence Estimation

ResearchDGX agent

arXiv:2510.10913v3 Announce Type: replace Abstract: Recent progress in large language models (LLMs) has enabled them to communicate their confidence in natural language, improving transparency and rel

Agent Capsules: Quality-Gated Granularity Control for Multi-Agent LLM Pipelines

SafetyDGX agent

arXiv:2605.00410v1 Announce Type: new Abstract: A multi-agent pipeline with N agents typically issues N LLM calls per run. Merging agents into fewer calls (compound execution) promises token savings,

AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go?

Model ReleasesDGX agent

arXiv:2605.00334v1 Announce Type: cross Abstract: Production agentic systems make many model calls per user request, and most of those calls are short, structured, and routine. This raises a practical

Agentic AI for Substance Use Education: Integrating Regulatory and Scientific Knowledge Sources

AgentsDGX agent

arXiv:2605.00383v1 Announce Type: new Abstract: The delivery of traditional substance education has remained problematic due to challenges in scalability, personalization, and the currency of informat

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs

Model ReleasesDGX agent

arXiv:2605.00539v1 Announce Type: new Abstract: Quantization is a key method for reducing the GPU memory requirement of training large language models (LLMs). Yet, current approaches are ineffective f

Alethia: A Foundational Encoder for Voice Deepfakes

Model ReleasesDGX agent

arXiv:2605.00251v1 Announce Type: cross Abstract: Existing voice deepfake detection and localization models rely heavily on representations extracted from speech foundation models (SFMs). However, dow

← Previous
1…8990919293…129
Next →