AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Safety

Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions

DGX agent

arXiv:2507.15692v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) provide new opportunities for blind and low vision (BLV) people to access visual information in their daily l

safetyarxiv-cs-cl
2 Jul 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Svarna: An Open Corpus Workbench for Modern Greek

DGX agent

arXiv:2607.00970v1 Announce Type: new Abstract: This paper introduces Svarna, a free, open-source, web-based corpus workbench for modern Greek. Svarna integrates five databases covering various regist

model-releasesarxiv-cs-cl
2 Jul 2026
Tutorials

The Course of News Events: A Comparison of Bottom-Up and Top-Down Approaches for Collecting Text-Based Data about Disasters

DGX agent

arXiv:2607.00849v1 Announce Type: new Abstract: News articles are an important source of information on disaster impacts and adaptation. A key methodological challenge in socio-environmental studies i

tutorialsarxiv-cs-cl
2 Jul 2026
Research

TRACE: State-Aware Query Processing over Temporal Evidence Graphs for Conversational Data

DGX agent

arXiv:2607.00339v1 Announce Type: new Abstract: Conversational data is increasingly used as a persistent source of user state for long-running assistants and AI agents. However, querying this data rem

researcharxiv-cs-cl
2 Jul 2026
Research

Understanding Large Language Models

DGX agent

arXiv:2607.01006v1 Announce Type: new Abstract: Large Language Models (LLMs) represent one of the most significant advances in AI and natural language processing in recent years. Still, many pressing

researcharxiv-cs-cl
2 Jul 2026
Safety

Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors

DGX agent

arXiv:2607.00447v1 Announce Type: new Abstract: Large language models often produce hallucinated answers that violate prompt-level constraints. A key diagnostic question is whether these failures refl

safetyarxiv-cs-cl
2 Jul 2026
Research

Watermarking for Proprietary Dataset Protection

DGX agent

arXiv:2607.00325v1 Announce Type: cross Abstract: A growing body of literature suggests that training data membership inference problems are fundamentally hard tasks in modern language modeling settin

researcharxiv-cs-cl
2 Jul 2026
Research

What Survives Into Context: A Diagnostic for Budget-Constrained Multi-Hop RAG and When Submodular Evidence Packing Improves It

DGX agent

arXiv:2607.00725v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) under a fixed reader-context budget forces a selection problem: of the evidence retrieved, only a fraction can be s

researcharxiv-cs-cl
2 Jul 2026
Research

When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers

DGX agent

arXiv:2607.00394v1 Announce Type: cross Abstract: LLM agents increasingly rely on retrieval buffers to store and reuse past experience, yet the cache management policies governing these buffers remain

researcharxiv-cs-cl
2 Jul 2026
Model Releases

YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese

DGX agent

arXiv:2607.00664v1 Announce Type: new Abstract: We propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese

model-releasesarxiv-cs-cl
2 Jul 2026
Model Releases

A Semantic-Layer-Mediated Agent for Natural Language to SQL over Heterogeneous Enterprise Databases

DGX agent

arXiv:2606.31041v1 Announce Type: new Abstract: Natural language-to-SQL (NL2SQL) over real-world enterprise databases remains significantly more challenging than on academic benchmarks. Enterprise sch

model-releasesarxiv-cs-cl
1 Jul 2026
Applications

Adapting Foundation ASR Models to Dysarthric Speech: A Case Study

DGX agent

arXiv:2606.31722v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems often perform poorly in dysarthric speech, limiting their usefulness to affected speakers in everyday communi

applicationsarxiv-cs-cl
1 Jul 2026
Model Releases

AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

DGX agent

arXiv:2606.31551v1 Announce Type: new Abstract: Training language models (LMs) remains a highly human-intensive process, even as frontier language model agents become increasingly capable at software

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Beyond Clean Text: Evaluating Encoder and Decoder Robustness for Bangla Event Detection in Noisy Text

DGX agent

arXiv:2606.30914v1 Announce Type: new Abstract: Event detection (ED) systems are typically evaluated on clean, curated text, leaving their robustness to real-world noise largely unexplored, particular

model-releasesarxiv-cs-cl
1 Jul 2026
Safety

BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding

DGX agent

arXiv:2606.31315v1 Announce Type: new Abstract: Speculative decoding accelerates inference by using a lightweight draft model to generate candidate tokens in parallel, and are then verified by the tar

safetyarxiv-cs-cl
1 Jul 2026
Model Releases

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams

DGX agent

arXiv:2606.22723v2 Announce Type: replace Abstract: Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended,

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Bridging Scientific Heritage: An Arabic--Russian Parallel Corpus and LLM Benchmark for Sustainable Knowledge Transfer

DGX agent

arXiv:2606.30943v1 Announce Type: new Abstract: Russian and Arabic are among the major languages of scientific communication. Language barriers impede the exchange of research results between these co

model-releasesarxiv-cs-cl
1 Jul 2026
Research

Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers

DGX agent

arXiv:2606.31779v1 Announce Type: cross Abstract: Language models typically reason via explicit chain-of-thought (CoT), generating intermediate steps token-by-token. Latent CoT offers an alternative:

researcharxiv-cs-cl
1 Jul 2026
Tutorials

Building a Multimodal Dataset of Academic Paper for Keyword Extraction

DGX agent

arXiv:2606.31069v1 Announce Type: new Abstract: Up to this point, keyword extraction task typically relies solely on textual data. Neglecting visual details and audio features from image and audio mod

tutorialsarxiv-cs-cl
1 Jul 2026
Model Releases

Building an ASR Solution for Training and Assessing Children's Reading

DGX agent

arXiv:2606.31508v1 Announce Type: new Abstract: Automatic speech recognition for children's reading remains underdeveloped for most African languages, including Bambara, despite its potential value fo

model-releasesarxiv-cs-cl
1 Jul 2026
Safety

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

DGX agent

arXiv:2606.31608v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A c

safetyarxiv-cs-cl
1 Jul 2026
Model Releases

Clinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question Answering

DGX agent

arXiv:2606.31432v1 Announce Type: new Abstract: Medical multiple-choice question answering requires parameter-efficient adaptation across heterogeneous knowledge domains and reasoning operations. A me

model-releasesarxiv-cs-cl
1 Jul 2026
Local Ai

CORTEX: Token-Level Hallucination Detection in RAG via Comparative Internal Representations

DGX agent

arXiv:2606.31033v1 Announce Type: new Abstract: In this paper, we propose CORTEX, a token-level hallucination detection method for Retrieval-Augmented Generation (RAG). In long-form RAG outputs, hallu

local-aiarxiv-cs-cl
1 Jul 2026
Agents

DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching

DGX agent

arXiv:2606.31980v1 Announce Type: new Abstract: Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a mul

agentsarxiv-cs-cl
1 Jul 2026
Model Releases

Explicit Fuzzy Logic in the Feed-Forward Layer: Self-Forgetting Quantifiers Discover Legible Grammatical-Licensing Detectors

DGX agent

arXiv:2606.31845v1 Announce Type: new Abstract: A transformer's feed-forward (FFN) sublayer materializes the distinctions attention gathers, yet gives no account of what it computes. In a parameter-ne

model-releasesarxiv-cs-cl
1 Jul 2026
Research

Exploring the relationship between team institutional composition and novelty in academic papers based on fine-grained knowledge entities

DGX agent

arXiv:2606.31058v1 Announce Type: new Abstract: The composition of author teams is an important factor influencing the novelty of academic papers. However, existing studies have paid limited attention

researcharxiv-cs-cl
1 Jul 2026
Model Releases

FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge

DGX agent

arXiv:2602.06625v2 Announce Type: replace Abstract: Existing LLM-as-a-Judge systems suffer from three fundamental limitations: limited adaptivity to task- and domain-specific evaluation criteria, syst

model-releasesarxiv-cs-cl
1 Jul 2026
Research

Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models

DGX agent

arXiv:2606.31511v1 Announce Type: cross Abstract: In deployment settings where retraining is infeasible, small frozen code models are routinely asked to repair a failed program after seeing their own

researcharxiv-cs-cl
1 Jul 2026
Research

Fork-Think with Confidence

DGX agent

arXiv:2606.31484v1 Announce Type: cross Abstract: Parallel thinking has enjoyed great success for boosting LLM performance on reasoning tasks without the need for any re-training. However, existing me

researcharxiv-cs-cl
1 Jul 2026
Safety

From Propositional to Perceptual Asymmetry: Extending Frictive Policy Optimization to Asymmetric Partial Information Dialogue

DGX agent

arXiv:2606.30973v1 Announce Type: new Abstract: Frictive Policy Optimization (FPO; Pustejovsky et al., 2025) treats friction in collaborative dialogue -- misalignment, misunderstanding, repair -- as a

safetyarxiv-cs-cl
1 Jul 2026
Model Releases

Generating consensus and dissent on massive discussion platforms with a semantic-vector model

DGX agent

arXiv:2601.13932v2 Announce Type: replace-cross Abstract: Reaching consensus on massive discussion networks is critical for reducing noise and achieving optimal collective outcomes. However, the natur

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Generative Skill Composition for LLM Agents

DGX agent

arXiv:2606.32025v1 Announce Type: new Abstract: Recent LLM agents benefit from skills for solving complex tasks. Skills encapsulate modular packages of procedural knowledge and instructions for perfor

model-releasesarxiv-cs-cl
1 Jul 2026
Research

Information Terra: A Narrative-Anchored Semantic-First Projection of Document Embeddings

DGX agent

arXiv:2606.30824v1 Announce Type: cross Abstract: We introduce Information Terra, a narrative-anchored semantic-first projection that places a document corpus on an Earth-like globe whose poles are tw

researcharxiv-cs-cl
1 Jul 2026
Safety

Linguistic Bias Mitigation for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck

DGX agent

arXiv:2606.31411v1 Announce Type: new Abstract: Rapid advancements in generative speech technology have compromised the reliability of voice biometrics. While current spoofing detectors excel when ass

safetyarxiv-cs-cl
1 Jul 2026
Research

Linguistic Distancing on Social Media: Indicators of Emotion Regulation Across Age Groups

DGX agent

arXiv:2606.30957v1 Announce Type: new Abstract: Managing our emotional responses to events is key to emotional well-being, a process referred to as emotion regulation in psychology. Previous work has

researcharxiv-cs-cl
1 Jul 2026
Model Releases

LLM-as-a-judge validity in physics assessment depends more on the task than the model

DGX agent

arXiv:2603.14732v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are increasingly considered for automated assessment and feedback, understanding when LLM marking is valid is

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

LOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype Alignment

DGX agent

arXiv:2606.31310v1 Announce Type: new Abstract: Fueled by increasing model scale and multimodal inputs, Multimodal Large Language Models (MLLMs) have emerged as a promising paradigm for Spoken Languag

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish

DGX agent

arXiv:2606.31947v1 Announce Type: new Abstract: State-of-the-art speech datasets predominantly focus on widely spoken languages, often overlooking low-resource languages such as Luxembourgish, which r

model-releasesarxiv-cs-cl
1 Jul 2026
Research

Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments

DGX agent

arXiv:2606.30987v1 Announce Type: new Abstract: Decision-makers routinely rely on expert judgments accompanied by written explanations, yet explanation quality is difficult to measure at scale. Foreca

researcharxiv-cs-cl
1 Jul 2026
Model Releases

Moral Safety in LLMs: Exposing Performative Compliance with Puzzled Cues

DGX agent

arXiv:2606.31644v1 Announce Type: new Abstract: As large language models take on morally consequential roles in healthcare, legal, and hiring contexts, we need to examine whether their ethical behavio

model-releasesarxiv-cs-cl
1 Jul 2026
Research

Multilingual Polarization Detection Using Transformer-Based Models with Class Weighting and Threshold Tuning

DGX agent

arXiv:2606.30857v1 Announce Type: new Abstract: This paper describes our submission to SemEval-2026 Task 9 on detecting multilingual, multicultural, and multievent online polarization. We address all

researcharxiv-cs-cl
1 Jul 2026
Research

Overview of the TalentCLEF 2026: Skill and Job Title Intelligence for Human Capital Management

DGX agent

arXiv:2606.31692v1 Announce Type: new Abstract: This paper presents an overview of the second edition of the TalentCLEF challenge, organized as a Lab at the Conference and Labs of the Evaluation Forum

researcharxiv-cs-cl
1 Jul 2026
Research

RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference

DGX agent

arXiv:2606.31519v1 Announce Type: cross Abstract: Long-context Large Language Model inference is severely bottlenecked by the massive Key-Value (KV) cache, yet existing sparse attention methods often

researcharxiv-cs-cl
1 Jul 2026
Research

Reference-Based Prosody and Rhythm Evaluation for Spoken Dialogue Systems

DGX agent

arXiv:2606.31055v1 Announce Type: new Abstract: Speech-to-speech (S2S) AI agents are advancing rapidly, yet evaluation lacks interpretable speech-native measures for conversational prosody and rhythm.

researcharxiv-cs-cl
1 Jul 2026
Safety

Rethinking On-policy Optimization for Query Augmentation

DGX agent

arXiv:2510.17139v3 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have led to a surge of interest in query augmentation for information retrieval (IR). Two main appro

safetyarxiv-cs-cl
1 Jul 2026
Model Releases

Review Residuals: Update-Conditioned Residual Gating for Transformers

DGX agent

arXiv:2606.31859v1 Announce Type: cross Abstract: Residual connections add every sublayer's proposed update with a fixed coefficient of one; the network never evaluates whether an update is reliable b

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Revising RVL-CDIP: Quantifying Errors and Test-Train Overlap

DGX agent

arXiv:2606.31446v1 Announce Type: new Abstract: RVL-CDIP is a popular dataset for benchmarking document classifiers. However, the dataset contains ample amounts of label errors as well as non-trivial

model-releasesarxiv-cs-cl
1 Jul 2026
Safety

Revocable Learned State via Process Sidecars

DGX agent

arXiv:2606.30788v1 Announce Type: cross Abstract: Language models are often adapted in stages: a public skill phase, a private memory phase, and a later safety phase that learns to refuse outputs tied

safetyarxiv-cs-cl
1 Jul 2026
← Previous
1…3536373839…161
Next →