AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
1 Jul 2026

AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

Model ReleasesDGX agent

arXiv:2606.31551v1 Announce Type: new Abstract: Training language models (LMs) remains a highly human-intensive process, even as frontier language model agents become increasingly capable at software

Beyond Clean Text: Evaluating Encoder and Decoder Robustness for Bangla Event Detection in Noisy Text

Model ReleasesDGX agent

arXiv:2606.30914v1 Announce Type: new Abstract: Event detection (ED) systems are typically evaluated on clean, curated text, leaving their robustness to real-world noise largely unexplored, particular

BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.31315v1 Announce Type: new Abstract: Speculative decoding accelerates inference by using a lightweight draft model to generate candidate tokens in parallel, and are then verified by the tar

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams

Model ReleasesDGX agent

arXiv:2606.22723v2 Announce Type: replace Abstract: Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended,

Bridging Scientific Heritage: An Arabic--Russian Parallel Corpus and LLM Benchmark for Sustainable Knowledge Transfer

Model ReleasesDGX agent

arXiv:2606.30943v1 Announce Type: new Abstract: Russian and Arabic are among the major languages of scientific communication. Language barriers impede the exchange of research results between these co

Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers

ResearchDGX agent

arXiv:2606.31779v1 Announce Type: cross Abstract: Language models typically reason via explicit chain-of-thought (CoT), generating intermediate steps token-by-token. Latent CoT offers an alternative:

Building a Multimodal Dataset of Academic Paper for Keyword Extraction

TutorialsDGX agent

arXiv:2606.31069v1 Announce Type: new Abstract: Up to this point, keyword extraction task typically relies solely on textual data. Neglecting visual details and audio features from image and audio mod

Building an ASR Solution for Training and Assessing Children's Reading

Model ReleasesDGX agent

arXiv:2606.31508v1 Announce Type: new Abstract: Automatic speech recognition for children's reading remains underdeveloped for most African languages, including Bambara, despite its potential value fo

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

SafetyDGX agent

arXiv:2606.31608v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A c

Clinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question Answering

Model ReleasesDGX agent

arXiv:2606.31432v1 Announce Type: new Abstract: Medical multiple-choice question answering requires parameter-efficient adaptation across heterogeneous knowledge domains and reasoning operations. A me

CORTEX: Token-Level Hallucination Detection in RAG via Comparative Internal Representations

Local AiDGX agent

arXiv:2606.31033v1 Announce Type: new Abstract: In this paper, we propose CORTEX, a token-level hallucination detection method for Retrieval-Augmented Generation (RAG). In long-form RAG outputs, hallu

DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching

AgentsDGX agent

arXiv:2606.31980v1 Announce Type: new Abstract: Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a mul

Explicit Fuzzy Logic in the Feed-Forward Layer: Self-Forgetting Quantifiers Discover Legible Grammatical-Licensing Detectors

Model ReleasesDGX agent

arXiv:2606.31845v1 Announce Type: new Abstract: A transformer's feed-forward (FFN) sublayer materializes the distinctions attention gathers, yet gives no account of what it computes. In a parameter-ne

Exploring the relationship between team institutional composition and novelty in academic papers based on fine-grained knowledge entities

ResearchDGX agent

arXiv:2606.31058v1 Announce Type: new Abstract: The composition of author teams is an important factor influencing the novelty of academic papers. However, existing studies have paid limited attention

FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge

Model ReleasesDGX agent

arXiv:2602.06625v2 Announce Type: replace Abstract: Existing LLM-as-a-Judge systems suffer from three fundamental limitations: limited adaptivity to task- and domain-specific evaluation criteria, syst

Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models

ResearchDGX agent

arXiv:2606.31511v1 Announce Type: cross Abstract: In deployment settings where retraining is infeasible, small frozen code models are routinely asked to repair a failed program after seeing their own

Fork-Think with Confidence

ResearchDGX agent

arXiv:2606.31484v1 Announce Type: cross Abstract: Parallel thinking has enjoyed great success for boosting LLM performance on reasoning tasks without the need for any re-training. However, existing me

From Propositional to Perceptual Asymmetry: Extending Frictive Policy Optimization to Asymmetric Partial Information Dialogue

SafetyDGX agent

arXiv:2606.30973v1 Announce Type: new Abstract: Frictive Policy Optimization (FPO; Pustejovsky et al., 2025) treats friction in collaborative dialogue -- misalignment, misunderstanding, repair -- as a

Generating consensus and dissent on massive discussion platforms with a semantic-vector model

Model ReleasesDGX agent

arXiv:2601.13932v2 Announce Type: replace-cross Abstract: Reaching consensus on massive discussion networks is critical for reducing noise and achieving optimal collective outcomes. However, the natur

Generative Skill Composition for LLM Agents

Model ReleasesDGX agent

arXiv:2606.32025v1 Announce Type: new Abstract: Recent LLM agents benefit from skills for solving complex tasks. Skills encapsulate modular packages of procedural knowledge and instructions for perfor

Information Terra: A Narrative-Anchored Semantic-First Projection of Document Embeddings

ResearchDGX agent

arXiv:2606.30824v1 Announce Type: cross Abstract: We introduce Information Terra, a narrative-anchored semantic-first projection that places a document corpus on an Earth-like globe whose poles are tw

Linguistic Bias Mitigation for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck

SafetyDGX agent

arXiv:2606.31411v1 Announce Type: new Abstract: Rapid advancements in generative speech technology have compromised the reliability of voice biometrics. While current spoofing detectors excel when ass

Linguistic Distancing on Social Media: Indicators of Emotion Regulation Across Age Groups

ResearchDGX agent

arXiv:2606.30957v1 Announce Type: new Abstract: Managing our emotional responses to events is key to emotional well-being, a process referred to as emotion regulation in psychology. Previous work has

LLM-as-a-judge validity in physics assessment depends more on the task than the model

Model ReleasesDGX agent

arXiv:2603.14732v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are increasingly considered for automated assessment and feedback, understanding when LLM marking is valid is

LOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype Alignment

Model ReleasesDGX agent

arXiv:2606.31310v1 Announce Type: new Abstract: Fueled by increasing model scale and multimodal inputs, Multimodal Large Language Models (MLLMs) have emerged as a promising paradigm for Spoken Languag

LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish

Model ReleasesDGX agent

arXiv:2606.31947v1 Announce Type: new Abstract: State-of-the-art speech datasets predominantly focus on widely spoken languages, often overlooking low-resource languages such as Luxembourgish, which r

Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments

ResearchDGX agent

arXiv:2606.30987v1 Announce Type: new Abstract: Decision-makers routinely rely on expert judgments accompanied by written explanations, yet explanation quality is difficult to measure at scale. Foreca

Moral Safety in LLMs: Exposing Performative Compliance with Puzzled Cues

Model ReleasesDGX agent

arXiv:2606.31644v1 Announce Type: new Abstract: As large language models take on morally consequential roles in healthcare, legal, and hiring contexts, we need to examine whether their ethical behavio

Multilingual Polarization Detection Using Transformer-Based Models with Class Weighting and Threshold Tuning

ResearchDGX agent

arXiv:2606.30857v1 Announce Type: new Abstract: This paper describes our submission to SemEval-2026 Task 9 on detecting multilingual, multicultural, and multievent online polarization. We address all

Overview of the TalentCLEF 2026: Skill and Job Title Intelligence for Human Capital Management

ResearchDGX agent

arXiv:2606.31692v1 Announce Type: new Abstract: This paper presents an overview of the second edition of the TalentCLEF challenge, organized as a Lab at the Conference and Labs of the Evaluation Forum

RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference

ResearchDGX agent

arXiv:2606.31519v1 Announce Type: cross Abstract: Long-context Large Language Model inference is severely bottlenecked by the massive Key-Value (KV) cache, yet existing sparse attention methods often

Reference-Based Prosody and Rhythm Evaluation for Spoken Dialogue Systems

ResearchDGX agent

arXiv:2606.31055v1 Announce Type: new Abstract: Speech-to-speech (S2S) AI agents are advancing rapidly, yet evaluation lacks interpretable speech-native measures for conversational prosody and rhythm.

Rethinking On-policy Optimization for Query Augmentation

SafetyDGX agent

arXiv:2510.17139v3 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have led to a surge of interest in query augmentation for information retrieval (IR). Two main appro

Review Residuals: Update-Conditioned Residual Gating for Transformers

Model ReleasesDGX agent

arXiv:2606.31859v1 Announce Type: cross Abstract: Residual connections add every sublayer's proposed update with a fixed coefficient of one; the network never evaluates whether an update is reliable b

Revising RVL-CDIP: Quantifying Errors and Test-Train Overlap

Model ReleasesDGX agent

arXiv:2606.31446v1 Announce Type: new Abstract: RVL-CDIP is a popular dataset for benchmarking document classifiers. However, the dataset contains ample amounts of label errors as well as non-trivial

Revocable Learned State via Process Sidecars

SafetyDGX agent

arXiv:2606.30788v1 Announce Type: cross Abstract: Language models are often adapted in stages: a public skill phase, a private memory phase, and a later safety phase that learns to refuse outputs tied

Robust Text Watermarking for Large Language Models via Dual Semantic Embeddings

SafetyDGX agent

arXiv:2606.31602v1 Announce Type: new Abstract: This work presents Dual-Embedding Watermarking (DEW), a semantic watermarking scheme for large language models (LLMs) that leverages contextual and toke

Scalable Behaviour Cloning on Browser Using via Skill Distillation

ApplicationsDGX agent

arXiv:2606.32014v1 Announce Type: new Abstract: Internet users collectively perform an enormous range of skilled work through web browsers, from software development and document editing to search, fo

SeKV: Resolution-Adaptive KV Cache with Hierarchical Semantic Memory for Long-Context LLM Inference

HardwareDGX agent

arXiv:2606.31145v1 Announce Type: new Abstract: Large language models increasingly operate over long contexts, where the KV cache becomes a dominant memory bottleneck: its size grows linearly with seq

SemRF: A Semantic Reference Frame for Residual-Stream Dynamics in Language Models

Model ReleasesDGX agent

arXiv:2606.32022v1 Announce Type: cross Abstract: Residual-stream analysis asks how language-model computation evolves across depth, but intermediate decoding requires comparable readout coordinates a

Signed-Permutation Coordinate Transport for RMSNorm Transformers

Model ReleasesDGX agent

arXiv:2606.31963v1 Announce Type: cross Abstract: Modern LLM workflows move coordinate-indexed objects across checkpoints: steering vectors, sparse autoencoders, top-k neuron sets, attribution lists,

SpikeLogBERT: Energy-Efficient Log Parsing Using Spiking Transformer Networks

Model ReleasesDGX agent

arXiv:2606.31781v1 Announce Type: cross Abstract: Log parsing is a fundamental step in automated log analysis, transforming raw system logs into structured event templates for downstream tasks such as

Symmetry in language statistics shapes the geometry of model representations

ResearchDGX agent

arXiv:2602.15029v3 Announce Type: replace-cross Abstract: The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a cir

TAG-DLM: Diffusion Language Models for Text-Attributed Graph Learning

Local AiDGX agent

arXiv:2606.31166v1 Announce Type: new Abstract: Text-attributed graphs (TAGs), where each node carries a natural language description, require models to jointly reason over text and graph topology. Ex

Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability

ResearchDGX agent

arXiv:2601.18778v3 Announce Type: replace-cross Abstract: RL methods for scaling large reasoning models stall on datasets with low initial success rates, and thus little training signal. We investigat

The Bidirectional Process Reward Model

Model ReleasesDGX agent

arXiv:2508.01682v3 Announce Type: replace Abstract: Process Reward Models (PRMs), which assign fine-grained scores to intermediate reasoning steps within a solution trajectory, have emerged as a promi

The Decomposition Is the Fingerprint: Per-Component Identity for Agent Skills

Model ReleasesDGX agent

arXiv:2606.31272v1 Announce Type: cross Abstract: AI agents increasingly acquire and execute skills at runtime: bundles of prompt instructions, executable code, and tool declarations fetched from mark

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action

Model ReleasesDGX agent

arXiv:2606.31916v1 Announce Type: new Abstract: Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in inc

Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition

ApplicationsDGX agent

arXiv:2606.31642v1 Announce Type: new Abstract: Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits pr

Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical Fallacies

Model ReleasesDGX agent

arXiv:2606.31039v1 Announce Type: new Abstract: Large Language Models (LLMs) exhibit strong semantic capabilities, yet their resilience to manipulative linguistic patterns such as logical fallacies re

Usage frequency and application variety of research methods in library and information science: Continuous investigation from 1991 to 2021

ResearchDGX agent

arXiv:2606.31081v1 Announce Type: cross Abstract: The present study analyzed over 26,000 research articles published between 1991 and 2021 in twenty-one major LIS (Library and Information Science) jou

Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale

AgentsDGX agent

arXiv:2606.30801v1 Announce Type: new Abstract: Personalization algorithms determine what content users encounter on online platforms. Auditing these systems is difficult because independent auditors

ViTL: Temporal Logic-Guided Zero-Shot Natural Language Navigation via Vision-Language Models

TutorialsDGX agent

arXiv:2606.30696v1 Announce Type: cross Abstract: Enabling robots to follow natural language commands to complete zero-shot long-horizon tasks remains challenging. It requires extracting implicit temp

What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR

Model ReleasesDGX agent

arXiv:2606.31112v1 Announce Type: new Abstract: ASR systems have been often reported to underperform on atypical speech. An often conflated compounding factor is the existence of two valid transcripti

What If We Allocate Test-Time Compute Adaptively?

Model ReleasesDGX agent

arXiv:2602.01070v5 Announce Type: replace Abstract: Test-time compute scaling allocates inference computation uniformly, uses fixed sampling strategies, and applies verification only for reranking. In

When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs

ResearchDGX agent

arXiv:2606.30814v1 Announce Type: new Abstract: Calibration evaluates whether a model confidence aligns with its empirical accuracy. Existing studies often compare the calibration of different large l

When the Database Fails: Prompting LLM Dialogue Agents for Safe Recovery in Task-Oriented Dialogue

Model ReleasesDGX agent

arXiv:2606.31307v1 Announce Type: new Abstract: Large language models used in task-oriented dialogue often produce fluent but unsafe responses when backend database calls fail, return empty results, o

30 Jun 2026

A Diagnostic Framework and Multi-Evaluator Audit of Evaluator-Driven Preference Dynamics in Self-Adapting LLM Agents

Model ReleasesDGX agent

arXiv:2606.29719v1 Announce Type: cross Abstract: Measurements of proprietary LLM evaluators can become invalid within weeks -- we document one case and provide the diagnostic framework to detect it.

A French OSCE Dialogue Dataset and Controllable Virtual Patient System for Clinical Training

ResearchDGX agent

arXiv:2606.28526v1 Announce Type: new Abstract: The clinical and communication skills of medical students are commonly assessed through Objective Structured Clinical Examinations (OSCEs), which consis

A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks

ApplicationsDGX agent

arXiv:2606.28531v1 Announce Type: cross Abstract: Automatically generated videos from scientific papers are increasingly used for education and research dissemination. However, existing evaluation met

← Previous
1…2829303132…129
Next →