AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
11 Aug 2026

Tree-of-Experience: Hierarchical Experience Management for Self-Evolving Agents

ResearchDGX agent

arXiv:2608.09044v1 Announce Type: new Abstract: Continual self-evolution requires LLM agents to transform environmental interactions into reliable and reusable experience. Existing methods typically r

Uncertainty-Aware Variational Reward Factorization via Probabilistic Preference Bases for LLM Personalization

SafetyDGX agent

arXiv:2604.00997v2 Announce Type: replace Abstract: Reward factorization personalizes large language models (LLMs) by decomposing rewards into shared basis functions and user-specific weights. Yet, ex

Universal or Language-Family-Specific Script Unification for Cross-Lingual Transfer? A Case Study on Turkic Languages

ApplicationsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.09356v1 Announce Type: new Abstract: Closely related languages written in different scripts expose little surface overlap to multilingual models, limiting cross-lingual transfer. We compare

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

Model ReleasesDGX agent

arXiv:2608.09209v1 Announce Type: new Abstract: Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic

UNSPECIFIC: General Constraint Synthesis for Breaking Copy-and-Paste Shortcut in LLM Instruction Following

Model ReleasesDGX agent

arXiv:2608.09154v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly expected to follow long lists of constraints in complex instructions, and synthesizing instructions from a

Unsure but Certain: Uncovering the Representation-Confidence Gap in Diffusion Language Models

ResearchDGX agent

arXiv:2608.08791v1 Announce Type: new Abstract: Diffusion language models use broad context to create text, suggesting they might handle input noise better than standard models. Testing reveals this i

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use

Model ReleasesDGX agent

arXiv:2608.08477v1 Announce Type: new Abstract: We present VectraYX-Vision-1B, a sub-2B vision-language model (VLM) for Spanish/LATAM cybersecurity imagery, coupling a frozen SigLIP-so400m encoder to

Verifiably grounded machine interpretation of lunar geology

Local AiDGX agent

arXiv:2608.09276v1 Announce Type: new Abstract: Planetary geology relies on historical, interpretive reasoning to reconstruct past events from diverse observations. Here, we present a step toward an a

VeriForge: Mitigating Latent Knowledge Gaps in Narrative Drafting via Mixed-Initiative Scaffolding

ResearchDGX agent

arXiv:2608.09698v1 Announce Type: cross Abstract: Great fiction earns its verisimilitude through precise details, from how a longsword is gripped to pierce armor gaps to why a bleeding corpse cannot y

Wisdom in Unity: The Role of Multilingual Training in Figurative Language Identification in Proverbs

ResearchDGX agent

arXiv:2608.08090v1 Announce Type: new Abstract: Although multilingual approaches to figurative language identification are not new, the shift beyond language homogeneous training data requires a clear

10 Aug 2026

A Practical Evaluation Method for Long-Form Simultaneous Speech-to-Speech Translation

SafetyDGX agent

arXiv:2606.15059v2 Announce Type: replace Abstract: Simultaneous speech-to-speech translation (SimulS2ST) enables real-time cross-lingual communication, but existing evaluation has focused largely on

AfriNLLB: Efficient Translation Models for African Languages

ResearchDGX agent

arXiv:2602.09373v2 Announce Type: replace Abstract: In this work, we present AfriNLLB, a series of lightweight models for efficient translation from and into African languages. AfriNLLB supports 15 la

An Exploratory Evaluation of LLM-Assisted Rewriting of Moderate-Complexity Financial Sentences for DisCoCat-Based Sentiment Analysis

Model ReleasesDGX agent

arXiv:2608.07439v1 Announce Type: new Abstract: Quantum natural language processing (QNLP) provides a grammar-aware framework for text modeling, and Distributional Compositional Categorical (DisCoCat)

Better Together: Quantifying the Benefits of AI-Assisted Recruitment

ResearchDGX agent

arXiv:2507.08029v2 Announce Type: replace Abstract: Hiring algorithms have mostly scored the materials recruiters already see. Large language models (LLMs) can instead generate new information about c

Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs

Model ReleasesDGX agent

arXiv:2608.06967v1 Announce Type: new Abstract: Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide

Confidence Estimation for Financial Vision-Language Models in Chart and Document Understanding

SafetyDGX agent

arXiv:2608.06532v1 Announce Type: new Abstract: LVLMs are increasingly used to read financial charts, tables, and documents, where a single misread figure can move a decision and the most authoritativ

Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models

ResearchDGX agent

arXiv:2608.06977v1 Announce Type: new Abstract: It is well established that large language models (LLMs) are sensitive to prompt framing, reflecting patterns in their training data or prior prompts. I

ConstructCIE: A Dataset for Extracting Causal Information from Construction Accident Narratives

ResearchDGX agent

arXiv:2608.06495v1 Announce Type: new Abstract: Construction accident narratives contain rich causal information, but the evidence is often implicit, long-span, and distributed. We introduce Construct

DAEP: Difficulty-Aware Evidence Planning for Medical Video Corpus Temporal Answer Grounding

Local AiDGX agent

arXiv:2608.06869v1 Announce Type: cross Abstract: We describe DAEP, team BIGC's submission to NLPCC 2026 Shared Task 1 Track 3: Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). T

Discovering Conceptual Metaphors Across Topics and Media Types

TutorialsDGX agent

arXiv:2608.06652v1 Announce Type: new Abstract: Conceptual metaphors guide our thinking and actions by allowing us to reason about more abstract experiences (e.g., paying taxes) in terms of more concr

Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation

Model ReleasesDGX agent

arXiv:2608.06718v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paraling

Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models?

ResearchDGX agent

arXiv:2608.07006v1 Announce Type: new Abstract: Visual retrieval-augmented generation (RAG) commonly expands the retrieved evidence set to improve answer-page coverage, implicitly assuming that all av

Don't `Well, Actually' Me Unless You Know What You're Talking About: Weak Presupposition Verification Degrades General QA Performance

TutorialsDGX agent

arXiv:2608.06539v1 Announce Type: new Abstract: False-presupposition QA (FPQA) tests LLMs on their ability to identify false presuppositions in questions and abstain or correct them rather than reinfo

From Test-Time Scaling to Reusable Memory: Measuring Crystallization in Text-to-SQL

ResearchDGX agent

arXiv:2608.07213v1 Announce Type: new Abstract: Test-time scaling can correct difficult text-to-SQL queries, but the extra computation is normally discarded after each answer. Systems increasingly ret

Gaze Behavior in Visual World Experiments Can be Modeled With Off-the-shelf Language-Vision Encoders

ResearchDGX agent

arXiv:2608.07282v1 Announce Type: new Abstract: The recent advances in neural language models have also spurred much work in computational psycholinguistics, asking whether neural LMs are also promisi

Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram Treebanks

ResearchDGX agent

arXiv:2608.07283v1 Announce Type: new Abstract: Grammar engineering requires expertise in linguistic formalism and computational implementation, especially in parallel grammar projects that balance cr

GRASP: Reinforcing Language Model Anonymizers with Group Relative Policy Optimization

Model ReleasesDGX agent

arXiv:2608.06526v1 Announce Type: new Abstract: Large language models can infer sensitive personal attributes, such as age, location, and occupation, from ordinary text, turning everyday writing into

HFS: Holistic Query-Aware Frame Selection for Efficient Video Understanding

ResearchDGX agent

arXiv:2512.11534v2 Announce Type: replace-cross Abstract: Key frame selection is essentially a set-level optimization problem: the quality of the selected subset depends on the interactions among fram

HNR-DAC: Hard-Negative Reranking and Distribution-Aligned Classification for Scientific Claim Verification

ResearchDGX agent

arXiv:2608.07204v1 Announce Type: new Abstract: Scientific claim verification over a cited paper requires predicting the claim--paper relation and identifying the paragraphs that justify that predicti

How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality

ResearchDGX agent

arXiv:2604.06756v2 Announce Type: replace Abstract: Large language models (LLMs) has been widely adopted as a scalable surrogate for human evaluation, yet such judges remain imperfect and susceptible

How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots

SafetyDGX agent

arXiv:2608.06898v1 Announce Type: cross Abstract: Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public l

Index SLM Technical Report

Model ReleasesDGX agent

arXiv:2607.09885v2 Announce Type: replace Abstract: We present Index-1.9B, a series of open small language models developed at Bilibili. The series comprises four models: Index-1.9B-Base, a foundation

Joint Optimization of Reasoning and Dual-Memory for Self-Learning Diagnostic Agent

AgentsDGX agent

arXiv:2604.07269v2 Announce Type: replace Abstract: Clinical expertise improves not only by acquiring medical knowledge, but by accumulating experience that yields reusable diagnostic patterns. Recent

Latent Fact-Checking: Detecting Misinformation through Activation Engineering

Model ReleasesDGX agent

arXiv:2608.06417v1 Announce Type: cross Abstract: The proliferation of misinformation online has driven demand for scalable detection systems. While most existing approaches rely on surface-level ling

Let's Unlearn Stereotypes Before Decision-Making: Assessing the Impact of Intrinsic Bias Mitigation on Downstream Fairness in LLMs

SafetyDGX agent

arXiv:2509.16462v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly used in high-stakes decision-making systems, where biased predictions can reinforce social and economi

LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering

Model ReleasesDGX agent

arXiv:2608.07370v1 Announce Type: new Abstract: Scientific literature is increasingly used as a knowledge source for language models, retrieval-augmented generation systems, and research assistants, b

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

Model ReleasesDGX agent

arXiv:2608.06867v1 Announce Type: new Abstract: No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment.

Lost in Interpolation: Why Predictive Feedback Fails in Diffusion Language Models

Model ReleasesDGX agent

arXiv:2608.06529v1 Announce Type: new Abstract: Soft-masking accelerates the convergence of Masked Diffusion Language Models (MDLMs). Existing formulations build this blend with linear interpolation (

Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand

ResearchDGX agent

arXiv:2608.06506v1 Announce Type: new Abstract: Language models are often evaluated as though capabilities demonstrated in English remain equally available when the same content is presented in other

Model Confidence Under Answer-Preserving Attacks: An Informativeness-Manipulability Frontier

Model ReleasesDGX agent

arXiv:2608.06571v1 Announce Type: cross Abstract: Deployed vision-language systems often gate their answers on confidence, making confidence robustness relevant to oversight. We study confidence reado

Modular TTT: Rethinking Test-Time Training as Composable Modules

Model ReleasesDGX agent

arXiv:2608.07110v1 Announce Type: cross Abstract: Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing

Model ReleasesDGX agent

arXiv:2608.06424v1 Announce Type: cross Abstract: Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re-synthesizing the entire utt

Multi-Perspective Triad Interaction Graph Neural Network for Cognitive Distortion Detection

SafetyDGX agent

arXiv:2608.06785v1 Announce Type: new Abstract: Cognitive distortion detection is a key task in computational mental health, yet existing approaches often overlook the psychological structure of disto

NTDH: Complex Reasoning for Comprehensive Affective Analysis

SafetyDGX agent

arXiv:2608.06425v1 Announce Type: new Abstract: Comprehensive affective analysis is challenging for two reasons: it spans heterogeneous prediction tasks with continuous, ordinal, and multi-label outpu

Pre-Inference Routing for Cost-Efficient Document Field Extraction

ResearchDGX agent

arXiv:2608.06607v1 Announce Type: new Abstract: Most document-extraction systems use a single model for all documents. This is simple but can be costly for easy cases and less effective for difficult

Quantization Damage Is Multiplicative, Not Additive

Model ReleasesDGX agent

arXiv:2608.06564v1 Announce Type: cross Abstract: Quantization is how large language models are actually deployed, and below four bits it is known to hurt. What nobody can say is which of the model's

Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models

ResearchDGX agent

arXiv:2608.06429v1 Announce Type: new Abstract: Interpretability methods for large language models (LLMs) describe internal state but do not directly test whether that state is causally sufficient to

Retrieval-Constrained Policy Optimization for Attack Technique Extraction from Cyber Threat Intelligence

Model ReleasesDGX agent

arXiv:2608.06778v1 Announce Type: cross Abstract: Mapping cyber threat intelligence (CTI) text to MITRE ATT&CK techniques is essential for structured threat analysis, yet manual annotation is costly a

Simple-OPD: Demystifying Warm-up for On-policy Distillation

Model ReleasesDGX agent

arXiv:2608.06802v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own rollouts with token-level supervision from teacher models, but its effectiveness can depend str

Skaling: Chinchilla's Exponents Meet Kaplan's Coupling

ResearchDGX agent

arXiv:2608.07222v1 Announce Type: new Abstract: Neural scaling laws are foundational for language model development, yet standard formulations systematically under- and overestimate loss at data-scarc

Stockmark-Nemotron-3-Nano-Omni-JapanDocReader: Structured Document Parsing via Capability Injection and Forgetting Control

Model ReleasesDGX agent

arXiv:2608.06758v1 Announce Type: new Abstract: We present Stockmark-Nemotron-3-Nano-Omni-JapanDocReader, a Japanese document understanding model built from Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16

Stoicheia: Character-Level Masked Diffusion for Ancient Greek Textual Restoration, Parsing, and Metrical Scansion

Model ReleasesDGX agent

arXiv:2608.07249v1 Announce Type: new Abstract: We introduce Stoicheia, a 405M-parameter character-level masked-diffusion encoder for Ancient Greek whose input factors into five aligned, independently

TA-RAG: Tone Awareness as a Design Imperative for Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2608.06672v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become a robust architecture for grounding large language models (LLMs) in trusted knowledge. However, standard

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents

SafetyDGX agent

arXiv:2608.06663v1 Announce Type: new Abstract: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning

AgentsDGX agent

arXiv:2608.07371v1 Announce Type: cross Abstract: Recent agentic reinforcement learning methods use hindsight to complement sparse outcome rewards. However, a completed rollout can yield many such sig

When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction

TutorialsDGX agent

arXiv:2505.16170v4 Announce Type: replace Abstract: We study the internal mechanisms that govern when LLMs choose to retract wrong answers, i.e., spontaneously and immediately acknowledge errors in th

Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models

TutorialsDGX agent

arXiv:2608.07261v1 Announce Type: new Abstract: Large language models (LLMs) can solve complex multi-hop problems yet exhibit puzzling failures on simple two-hop queries: although a model may correctl

7 Aug 2026

An Early Warning of Emerging Biosecurity Risks in Frontier LLMs

Model ReleasesDGX agent

arXiv:2607.18056v2 Announce Type: replace Abstract: Frontier large language models (LLMs) are increasingly integrated into scientific workflows, yet their growing biological capabilities may outpace c

Analysis of Numerical Localisation in LLM Translations

Local AiDGX agent

arXiv:2608.05232v1 Announce Type: new Abstract: The work of Tang et. al. (2025) on numerical translation is extended by analysing the capability of five large language models (LLMs) for the localisati

Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents

Model ReleasesDGX agent

arXiv:2608.06312v1 Announce Type: new Abstract: Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficie

← Previous
123456…128
Next →