AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
2 Jul 2026

AGC-Bench: Measuring Artificial General Creativity

Model ReleasesDGX agent

arXiv:2607.01152v1 Announce Type: new Abstract: Creativity research has debated whether creativity is domain-specific (e.g., visual, writing, science), and if it is psychometrically separable from gen

ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs

ResearchDGX agent

arXiv:2607.00171v1 Announce Type: new Abstract: Text embeddings are standard for semantic similarity tasks, yet their evaluation remains an open challenge. Current benchmarks are static, cover only a

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.00139v1 Announce Type: new Abstract: The cost of human expert evaluation is a principal bottleneck to deploying language models in specialized, high-stakes domains. This is particularly acu

Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents

Model ReleasesDGX agent

arXiv:2607.00895v1 Announce Type: new Abstract: Hallucination detection for retrieval-augmented generation (RAG) is usually evaluated on natural-language document evidence. However, grounded generatio

Beyond Perplexity: A Behavioral Evaluation Framework for Deployment-Memory Claims in LLM Test-Time Training

Local AiDGX agent

arXiv:2607.00368v1 Announce Type: new Abstract: Large language model test-time training (TTT) is often evaluated through local proxy metrics: models are updated on recent tokens, retrieved context, ta

Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking

Model ReleasesDGX agent

arXiv:2607.01103v1 Announce Type: new Abstract: Open-response evaluation provides stronger clinical validity than multiple-choice benchmarks but creates a scoring bottleneck that motivates automated L

CogTax: A Four-Level Cognitive Taxonomy for Command-Line Computing Education

TutorialsDGX agent

arXiv:2607.00140v1 Announce Type: cross Abstract: As computing education expands beyond traditional programming into operational domains such as systems administration and command-line environments, e

Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates

AgentsDGX agent

arXiv:2607.01047v1 Announce Type: new Abstract: Complexity and interpretability rarely coincide: systems rich enough for complex behaviours to emerge are usually too opaque to question, while transpar

Disentangling Speaker and Language Effects in Cross-Lingual Speaker Verification for Iberian Languages

ResearchDGX agent

arXiv:2607.01161v1 Announce Type: cross Abstract: Cross-lingual speaker verification (SV) systems typically exhibit performance degradation when enrollment and test utterances are spoken in different

'Don't Say It!': Constraints, Compliance, and Communication when Language Models Play Taboo

ResearchDGX agent

arXiv:2607.00601v1 Announce Type: new Abstract: The game of Taboo requires describing a target word without using a set of forbidden words, so that other players can guess it. This deceptively simple

Dual-Confidence Contrastive Decoding for Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2607.00570v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) increasingly requires models to answer questions from multiple retrieved documents, where only some sources are rel

Dynamic Bidirectional Pattern Memory: A Production-Scale Empirical Characterisation of Inference-Time Gating in Clinical NLP

Model ReleasesDGX agent

arXiv:2607.00870v1 Announce Type: new Abstract: We study inference-time pattern-memory gating in a production-scale clinical natural language processing (NLP) pipeline. The pipeline pairs a generator

Efficient Multilingual Reasoning Transfer via Progressive Code-Switching

ResearchDGX agent

arXiv:2607.00485v1 Announce Type: new Abstract: Large reasoning models (LRMs) have achieved strong reasoning capabilities in English, yet their performance degrades significantly when required to reas

EPC: A Standardized Protocol for Measuring Evaluator Preference Dynamics in LLM Agent Systems

Model ReleasesDGX agent

arXiv:2607.00297v1 Announce Type: cross Abstract: When LLM agents use evaluator feedback to adapt their behavior in closed loops, evaluator biases propagate through the agent's strategy distribution -

Evidence-Supported Credit Risk Report Generation Using News-Centric Financial Knowledge Graphs

ApplicationsDGX agent

arXiv:2607.01023v1 Announce Type: new Abstract: Financial markets evolve in response to real-world events reported in news, yet these drivers often remain implicit in text. To better explain market dy

ext{Log}_ext{b}Quant: Quantizing Language Models in Logarithmic Space

Model ReleasesDGX agent

arXiv:2607.01127v1 Announce Type: new Abstract: Quantization has become an invaluable tool to reduce memory requirements and inference speed of modern language models, in particular to make them avail

From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape

SafetyDGX agent

arXiv:2606.08625v2 Announce Type: replace Abstract: As Large Language Models (LLMs) advance toward open-ended autonomous agents, the mechanisms used to evaluate and guide their behavior must evolve ac

Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization

AgentsDGX agent

arXiv:2601.04424v2 Announce Type: replace Abstract: Large language models (LLMs) now support contexts of up to 1M tokens, but their strengths and weaknesses on complex long-context tasks remain unclea

GPTKB v1.5: A Massive Knowledge Base for Exploring Factual LLM Knowledge

Model ReleasesDGX agent

arXiv:2507.05740v2 Announce Type: replace Abstract: Language models are powerful artifacts, yet their factual knowledge is still poorly understood, and inaccessible to ad-hoc browsing and scalable sta

Graded strength of comparative illusions is explained by Bayesian inference

ResearchDGX agent

arXiv:2511.14642v2 Announce Type: replace Abstract: Like visual processing, language processing is susceptible to illusions in which people systematically misperceive stimuli. In one such case--the co

How Do We Engage with Other Disciplines? A Framework to Study Meaningful Interdisciplinary Discourse in Scholarly Publications

ResearchDGX agent

arXiv:2601.17020v2 Announce Type: replace-cross Abstract: With the rising popularity of interdisciplinary work and increasing institutional incentives in this direction, there is a growing need to und

How Ethos and Pathos Appeals Resonate in Reader Interpretations of Social Media Messages

ResearchDGX agent

arXiv:2607.00873v1 Announce Type: new Abstract: Rhetorical strategies and their influence on audiences are often studied through social media posts and comments. However, this focus overlooks the univ

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting

Model ReleasesDGX agent

arXiv:2607.00159v1 Announce Type: new Abstract: Knowledge-Based Visual Question Answering (KB-VQA) aims to evaluate whether Visual Language Models (VLMs) can retrieve, ground, and reason over external

Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training

Model ReleasesDGX agent

arXiv:2607.01232v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central component of post-training large language models (LLMs), yet little is understood about how RL adapta

Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking

ResearchDGX agent

arXiv:2607.00482v1 Announce Type: new Abstract: Reasoning language models frequently overthink: generating extended chains of behaviors such as hedging, approach abandonment, and self contradiction th

KnowledgeDebugger -- an Exploration Tool for Knowledge Localization and Editing in Transformers

ResearchDGX agent

arXiv:2607.01000v1 Announce Type: new Abstract: Recent research has increasingly focused on understanding how Transformers store and process knowledge, as well as how this knowledge can be edited. Res

Low Perplexity is Repetition: A One-Dimensional Self-Conditioning Attractor in Continuous Diffusion LMs

ResearchDGX agent

arXiv:2607.00588v1 Announce Type: new Abstract: Continuous diffusion language models such as ELF report record-low generative perplexity (Gen-PPL). We find a catch: these models repeat far more than h

LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data

Model ReleasesDGX agent

arXiv:2510.24434v3 Announce Type: replace Abstract: The effectiveness of instruction-tuned Large Language Models (LLMs) is often limited in low-resource linguistic settings due to a lack of high-quali

LV-ROVER: Multi-Stream Tesseract Voting for Maltese Paragraph OCR

Model ReleasesDGX agent

arXiv:2607.00250v1 Announce Type: new Abstract: Maltese has decent text corpora and pretrained language models, but, like many languages outside the handful with large OCR benchmarks, only a single kn

Message Passing Enables Efficient Reasoning

ResearchDGX agent

arXiv:2607.01077v1 Announce Type: new Abstract: While inference-time scaling has improved the reasoning abilities of large language models (LLMs), the need to generate long chains-of-thought (CoTs) is

MetaHOPE: A Metaphor-Oriented Evaluation Framework for Analysing MT and LLM Translation Errors

ResearchDGX agent

arXiv:2607.00848v1 Announce Type: new Abstract: In this opinion paper, we propose MetaHOPE, an error severity-aware annotation framework for evaluating metaphor translations. Metaphors present challen

MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules

Model ReleasesDGX agent

arXiv:2607.00464v1 Announce Type: cross Abstract: Current molecular generation benchmarks emphasize task complexity, molecule novelty, and property alignment; they largely overlook a critical concern:

MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark

Model ReleasesDGX agent

arXiv:2607.00724v1 Announce Type: new Abstract: Multilingual fluency often invites a stronger assumption: a model that can speak a user's language must also understand the culture encoded by that lang

Multi-Turn Agentic Scientific Literature Search via Workflow Induction

AgentsDGX agent

arXiv:2607.00597v1 Announce Type: new Abstract: Scientific literature search often requires more than retrieving papers from a single query: users' intents are underspecified, preference-dependent, an

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

Model ReleasesDGX agent

arXiv:2607.00890v1 Announce Type: new Abstract: Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

SafetyDGX agent

arXiv:2510.24636v3 Announce Type: replace Abstract: Reward models (RMs) have become essential for aligning large language models (LLMs), serving as scalable proxies for human evaluation in both traini

Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions

ResearchDGX agent

arXiv:2607.00937v1 Announce Type: new Abstract: Persona-driven generations (PDGs) have seen prolific use in research and industry applications, where a large language model (LLM) takes on a 'persona'

Quantifying the Affective Gap: A Zero-Shot Evaluation of LLMs on Fine-Grained Emotion Taxonomies

Model ReleasesDGX agent

arXiv:2607.00968v1 Announce Type: new Abstract: Emotion recognition in natural language is a foundational challenge in affective computing, with critical implications for human-computer interaction, m

QuasiMoTTo: Quasi-Monte Carlo Test-Time Scaling

SafetyDGX agent

arXiv:2607.01179v1 Announce Type: cross Abstract: Scaling inference compute, by generating many parallel attempts per problem, is a costly but reliable lever for improving language model capabilities.

Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination

Local AiDGX agent

arXiv:2607.00158v1 Announce Type: new Abstract: Hallucination remains one of the central obstacles to deploying medical LLMs. Yet, even when hallucination can be detected, it is still unclear whether

Rosetta: Composable Native Multimodal Pretraining

ResearchDGX agent

arXiv:2607.00293v1 Announce Type: cross Abstract: Achieving true artificial general intelligence requires foundation models capable of integrating new modalities without forgetting prior knowledge. Ho

Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine

SafetyDGX agent

arXiv:2607.00576v1 Announce Type: new Abstract: Multi-image content has become an increasingly prevalent form of visual communication in social media, giving rise to a new safety issue, multi-image im

Selective Test-Time Debiasing for CLIP via Reward Gating

SafetyDGX agent

arXiv:2607.00423v1 Announce Type: new Abstract: Vision language models (VLMs) demonstrate strong zero-shot performance, but often perpetuate social stereotypes in person-centric queries, yielding skew

SemEval-2026 Task 9: Detecting Multilingual, Multicultural and Multievent Online Polarization

ResearchDGX agent

arXiv:2604.06817v2 Announce Type: replace Abstract: We present SemEval-2026 Task 9, a shared task on online polarization detection, covering 22 languages and comprising over 110K annotated instances.

SlowBA: An efficiency backdoor attack towards VLM-based GUI agents

AgentsDGX agent

arXiv:2603.08316v3 Announce Type: replace-cross Abstract: Modern vision-language-model (VLM) based graphical user interface (GUI) agents are expected not only to execute actions accurately but also to

Speech Playground: An Interactive Tool for Speech Analysis and Comparison

SafetyDGX agent

arXiv:2607.00418v1 Announce Type: new Abstract: This paper presents Speech Playground, an interactive speech visualization and comparison tool. While existing tools such as Praat are excellent, it can

StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning

Model ReleasesDGX agent

arXiv:2607.00465v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) rely extensively on Visual Instruction Tuning (VIT) to elicit their multimodal reasoning capabilities. However, w

Structural Pattern Mining in Inka Khipus: Unsupervised Clustering, Provenance Classification, and a Computational Validation of the Santa Valley Match

ResearchDGX agent

arXiv:2607.00185v1 Announce Type: new Abstract: Khipus--knotted cord devices--were the primary recording medium of the Inka Empire (c. 1400-1532 CE), yet their system remains undeciphered. We present

Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions

SafetyDGX agent

arXiv:2507.15692v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) provide new opportunities for blind and low vision (BLV) people to access visual information in their daily l

Svarna: An Open Corpus Workbench for Modern Greek

Model ReleasesDGX agent

arXiv:2607.00970v1 Announce Type: new Abstract: This paper introduces Svarna, a free, open-source, web-based corpus workbench for modern Greek. Svarna integrates five databases covering various regist

The Course of News Events: A Comparison of Bottom-Up and Top-Down Approaches for Collecting Text-Based Data about Disasters

TutorialsDGX agent

arXiv:2607.00849v1 Announce Type: new Abstract: News articles are an important source of information on disaster impacts and adaptation. A key methodological challenge in socio-environmental studies i

TRACE: State-Aware Query Processing over Temporal Evidence Graphs for Conversational Data

ResearchDGX agent

arXiv:2607.00339v1 Announce Type: new Abstract: Conversational data is increasingly used as a persistent source of user state for long-running assistants and AI agents. However, querying this data rem

Understanding Large Language Models

ResearchDGX agent

arXiv:2607.01006v1 Announce Type: new Abstract: Large Language Models (LLMs) represent one of the most significant advances in AI and natural language processing in recent years. Still, many pressing

Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors

SafetyDGX agent

arXiv:2607.00447v1 Announce Type: new Abstract: Large language models often produce hallucinated answers that violate prompt-level constraints. A key diagnostic question is whether these failures refl

Watermarking for Proprietary Dataset Protection

ResearchDGX agent

arXiv:2607.00325v1 Announce Type: cross Abstract: A growing body of literature suggests that training data membership inference problems are fundamentally hard tasks in modern language modeling settin

What Survives Into Context: A Diagnostic for Budget-Constrained Multi-Hop RAG and When Submodular Evidence Packing Improves It

ResearchDGX agent

arXiv:2607.00725v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) under a fixed reader-context budget forces a selection problem: of the evidence retrieved, only a fraction can be s

When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers

ResearchDGX agent

arXiv:2607.00394v1 Announce Type: cross Abstract: LLM agents increasingly rely on retrieval buffers to store and reuse past experience, yet the cache management policies governing these buffers remain

YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese

Model ReleasesDGX agent

arXiv:2607.00664v1 Announce Type: new Abstract: We propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese

1 Jul 2026

A Semantic-Layer-Mediated Agent for Natural Language to SQL over Heterogeneous Enterprise Databases

Model ReleasesDGX agent

arXiv:2606.31041v1 Announce Type: new Abstract: Natural language-to-SQL (NL2SQL) over real-world enterprise databases remains significantly more challenging than on academic benchmarks. Enterprise sch

Adapting Foundation ASR Models to Dysarthric Speech: A Case Study

ApplicationsDGX agent

arXiv:2606.31722v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems often perform poorly in dysarthric speech, limiting their usefulness to affected speakers in everyday communi

← Previous
1…2728293031…129
Next →