AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlog
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Model Releases

NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap

DGX agent

arXiv:2608.04397v1 Announce Type: new Abstract: We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 p

model-releasesarxiv-cs-cl
6 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation

DGX agent

arXiv:2606.17188v3 Announce Type: replace-cross Abstract: Current multilingual evaluations for Vision-Language Models (VLMs) assume a one-to-one mapping between language and orthography, overlooking b

model-releasesarxiv-cs-cl
6 Aug 2026
Safety

ODRA: Synthesizing Cognitive Behavioral Therapy Sessions with Structured Chain-Of-Thought and Dynamic Patient Resistance

DGX agent

arXiv:2608.04524v1 Announce Type: new Abstract: Synthetic generation of Cognitive Behavioral Therapy (CBT) sessions is challenged by two competing demands: adhering to strict therapeutic structure whi

safetyarxiv-cs-cl
6 Aug 2026
Safety

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning

DGX agent

arXiv:2608.05080v1 Announce Type: cross Abstract: Critic-free group-based reinforcement learning has become a scalable approach for post-training large language models. However, most existing methods

safetyarxiv-cs-cl
6 Aug 2026
Model Releases

Predict, Then Retrieve: Cross-Instance Future-State Retrieval from Video Prefixes

DGX agent

arXiv:2608.04426v1 Announce Type: cross Abstract: We introduce Predictive State Retrieval (PSR), a task in which a model observes a short video prefix and a temporal question about an object's future

model-releasesarxiv-cs-cl
6 Aug 2026
Research

Preverbal Uninflected and Underived Roots in Mapudungun. Wuno and Its Implications

DGX agent

arXiv:2608.04869v1 Announce Type: new Abstract: This study examines the grammatical status of preverbal uninflected and underived roots in Mapudungun, with particular focus on wuno 'return/re-'. Throu

researcharxiv-cs-cl
6 Aug 2026
Research

Reachability in 3-VAS

DGX agent

arXiv:2608.04786v1 Announce Type: new Abstract: We settle the exact complexity of the reachability problem in (stateless) vector addition systems (VAS) in fixed low dimension. In dimensions 2-4 it has

researcharxiv-cs-cl
6 Aug 2026
Model Releases

Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos

DGX agent

arXiv:2608.04939v1 Announce Type: new Abstract: Social media videos often communicate meanings that go beyond their visible actions, captions, or speech. A mundane clip may become humorous, ironic, or

model-releasesarxiv-cs-cl
6 Aug 2026
Research

Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training

DGX agent

arXiv:2608.05148v1 Announce Type: new Abstract: Procedural generators produce useful verifiable reasoning problems at scale, but have received less attention as data for completion-supervised fine-tun

researcharxiv-cs-cl
6 Aug 2026
Research

Relational Response Fields: A General Theory of Black-Box LLM Response Consistency and Recovery

DGX agent

arXiv:2608.04552v1 Announce Type: new Abstract: Black-box language-model reliability is commonly pursued by sampling, prompting, voting, verifying, or iteratively revising individual answers. We ask a

researcharxiv-cs-cl
6 Aug 2026
Model Releases

Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression

DGX agent

arXiv:2608.04569v1 Announce Type: new Abstract: Hard prompt compression reduces long-context inference cost by independently scoring tokens, sentences, or chunks and retaining the highest-scoring unit

model-releasesarxiv-cs-cl
6 Aug 2026
Research

Representing Visual Evidence for Item Difficulty Prediction: Visual Textualization and Image-Native Modeling

DGX agent

arXiv:2608.04554v1 Announce Type: new Abstract: Predicting item difficulty from content can provide an initial estimate for newly developed questions before sufficient student responses are available.

researcharxiv-cs-cl
6 Aug 2026
Model Releases

RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

DGX agent

arXiv:2608.04514v1 Announce Type: new Abstract: Background: Respiratory specialty care requires multimodal interpretation, longitudinal risk assessment, guideline-concordant intervention, and whole-co

model-releasesarxiv-cs-cl
6 Aug 2026
Research

Right Reset: Chunking by Prefix Removal

DGX agent

arXiv:2608.04330v1 Announce Type: new Abstract: Removing the left context from a causal language model reveals a useful kind of boundary: an edge where the model processes the same right-hand tokens w

researcharxiv-cs-cl
6 Aug 2026
Research

RingSQL: Schema-Independent Synthetic Data Generation for Text-to-SQL Reinforcement Learning

DGX agent

arXiv:2601.05451v2 Announce Type: replace-cross Abstract: Recent advances in text-to-SQL have been driven by larger models, better datasets, and new training methods like RLVR. However, progress remai

researcharxiv-cs-cl
6 Aug 2026
Model Releases

Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications?

DGX agent

arXiv:2608.05097v1 Announce Type: new Abstract: Reasoning about necessity and possibility depends on assumptions about accessibility between worlds and about which objects exist at each one. The same

model-releasesarxiv-cs-cl
6 Aug 2026
Research

Searching for Sound-Meaning Collisions: Graph-Based Affordance Retrieval and Multi-Evaluator Ranking for Pun Translation at CLEF 2026 JOKER Task 2

DGX agent

arXiv:2608.04299v1 Announce Type: new Abstract: Fifteen years ago, Low proposed that pun translators should stop searching for equivalent words and instead search for new points of contact between sou

researcharxiv-cs-cl
6 Aug 2026
Model Releases

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

DGX agent

arXiv:2608.04244v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks

model-releasesarxiv-cs-cl
6 Aug 2026
Research

Simile Understanding in Text-to-Image Models: An Evaluation Framework

DGX agent

arXiv:2608.04750v1 Announce Type: cross Abstract: Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce

researcharxiv-cs-cl
6 Aug 2026
Model Releases

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

DGX agent

arXiv:2608.04828v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools

model-releasesarxiv-cs-cl
6 Aug 2026
Safety

Social Pressure Breaks Majority Voting in LLM Safety Panels

DGX agent

arXiv:2608.04415v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct

safetyarxiv-cs-cl
6 Aug 2026
Model Releases

SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization

DGX agent

arXiv:2608.04084v1 Announce Type: cross Abstract: Mixture-of-experts (MoE) networks pursue specialization through learned routers, gates, and load-balancing losses, yet at matched total-parameter budg

model-releasesarxiv-cs-cl
6 Aug 2026
Safety

SpecRoll: Fast-Slow Verifier-Feedback Adaptation for Speculative Reinforcement Learning Rollouts

DGX agent

arXiv:2608.04962v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training improves the reasoning capabilities of large language models, but autoregressive rollout generation remains

safetyarxiv-cs-cl
6 Aug 2026
Agents

Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models

DGX agent

arXiv:2608.05126v1 Announce Type: new Abstract: Spoken Language Understanding (SLU) is the core component of task-oriented dialogue systems and a pivotal link in achieving seamless human-agent interac

agentsarxiv-cs-cl
6 Aug 2026
Agents

State2State: Environment-Derived Mid-Training for LLM Agents

DGX agent

arXiv:2608.04934v1 Announce Type: new Abstract: Training LLM agents commonly relies on supervised fine-tuning from expert trajectories or online reinforcement learning over human-specified tasks with

agentsarxiv-cs-cl
6 Aug 2026
Model Releases

Strengthening Target-Language Features: SAE-Based Steering for Multilingual Inference

DGX agent

arXiv:2608.04904v1 Announce Type: new Abstract: Multilingual large language models exhibit substantial performance differences across languages, while existing adaptation methods often require paramet

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation

DGX agent

arXiv:2608.04567v1 Announce Type: new Abstract: Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language p

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

Test, then Route: How Language Models Execute In-Context Conditional Rules Across Models and Languages

DGX agent

arXiv:2608.04183v1 Announce Type: new Abstract: When a language model follows an in-context conditional rule such as 'if P(x) then A else B,' does it assemble a runtime circuit with one module that te

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

The Calibration Floor: Format Repair Can Masquerade as Self-Correction at Small-to-Mid Scale

DGX agent

arXiv:2608.04355v1 Announce Type: new Abstract: Accuracy changes after language-model self-revision are usually interpreted as changes in reasoning. We show this can fail at the answer-extraction boun

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

The Evaluator Is Part of the Experiment: Measuring Open-Ended LLM Conformity

DGX agent

arXiv:2608.04463v1 Announce Type: new Abstract: Prior work on LLM conformity largely measures discrete answer flips under verifiable labels. Open-ended revisions require a different measurement strate

model-releasesarxiv-cs-cl
6 Aug 2026
Safety

The Fairness Collapse Phenomenon: Bias Amplification in Language Models Trained on Synthetic Data

DGX agent

arXiv:2608.04268v1 Announce Type: new Abstract: Generative models trained on artificially generated data have been shown to exhibit model collapse, resulting in significant performance degradation. As

safetyarxiv-cs-cl
6 Aug 2026
Research

The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

DGX agent

arXiv:2608.04570v1 Announce Type: new Abstract: Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inferenc

researcharxiv-cs-cl
6 Aug 2026
Agents

TopoChunker: Topology-Aware Agentic Document Chunking Framework

DGX agent

arXiv:2603.18409v2 Announce Type: replace Abstract: Current document chunking methods for Retrieval-Augmented Generation (RAG) typically linearize text. This forced linearization strips away intrinsic

agentsarxiv-cs-cl
6 Aug 2026
Model Releases

Toward Federated Large Language Models in Medicine: A Parameter-Efficient Framework for Privacy-Preserving, Multi-Institutional Adaptation

DGX agent

arXiv:2601.22124v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly adapted for medical applications, but most are trained using data from a single institution because pr

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

DGX agent

arXiv:2608.05139v1 Announce Type: new Abstract: Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivat

model-releasesarxiv-cs-cl
6 Aug 2026
Research

Towards End-to-End Multilingual Metaphor Processing: Integrating Detection, Translation, and Evaluation

DGX agent

arXiv:2608.04260v1 Announce Type: new Abstract: Metaphorical language remains a major challenge for multilingual natural language processing because successful interpretation and translation require r

researcharxiv-cs-cl
6 Aug 2026
Local Ai

Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs

DGX agent

arXiv:2608.04759v1 Announce Type: cross Abstract: Although Multimodal Large Language Models (MLLMs) have made substantial progress, their spatial reasoning may still produce intermediate judgments inc

local-aiarxiv-cs-cl
6 Aug 2026
Model Releases

Transfer Learning for Named Entity Recognition of Classical Latin through LLM Prompting

DGX agent

arXiv:2608.04015v1 Announce Type: new Abstract: With the increase in digitized resources of Classical Latin texts and modern breakthroughs of Large Language Models (LLMs), I contribute to ancient lang

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

DGX agent

arXiv:2607.19322v2 Announce Type: replace Abstract: Rubric-based evaluation of open-ended generation faces a fundamental tension between expressiveness and reliability. Authoring a faithful rubric req

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

UG-UMRE: Uncertainty-Guided Modality Augmentation and Distributional Calibration for Unified Multimodal Relation Extraction

DGX agent

arXiv:2608.04949v1 Announce Type: cross Abstract: Unified Multimodal Relation Extraction (UMRE) aims to identify intra-modal and cross-modal relations between textual entities and visual objects. Howe

model-releasesarxiv-cs-cl
6 Aug 2026
Tutorials

Unforgettable Generalization in Language Models

DGX agent

arXiv:2409.02228v2 Announce Type: replace-cross Abstract: When language models (LMs) are trained to forget (or 'unlearn'') a skill, how precisely does their behavior change? We study the behavior of t

tutorialsarxiv-cs-cl
6 Aug 2026
Safety

When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents

DGX agent

arXiv:2608.04574v1 Announce Type: new Abstract: Memory-augmented VLM agents act on persistent spatial knowledge, yet that knowledge silently goes stale as the environment changes. We ask what happens

safetyarxiv-cs-cl
6 Aug 2026
Safety

When Modalities Remember: Continual Learning for Multimodal Knowledge Graphs

DGX agent

arXiv:2604.02778v2 Announce Type: replace Abstract: Real-world multimodal knowledge graphs (MMKGs) are dynamic, with new entities, relations, and multimodal knowledge emerging over time. Existing cont

safetyarxiv-cs-cl
6 Aug 2026
Research

When More Becomes Less: Position-Dependent Repetition Effects in Language Models

DGX agent

arXiv:2608.04021v1 Announce Type: new Abstract: Cloze-style probes that vary how often a target token appears implicitly assume that more copies of a target affect prediction the same way regardless o

researcharxiv-cs-cl
6 Aug 2026
Research

A machine-readable catalogue of the Tsiolkovsky papers (fond 555, Archive of the Russian Academy of Sciences), and a way to measure how well its handwriting can be read

DGX agent

arXiv:2608.03617v1 Announce Type: new Abstract: The personal archive of Konstantin Tsiolkovsky (1857-1935) is held as fond 555 of the Archive of the Russian Academy of Sciences. The archive scanned th

researcharxiv-cs-cl
5 Aug 2026
Hardware

AcceptMoE: Commitment-Weighted Self-Sizing Verifier Expert Sets for Efficient MoE Speculative Decoding

DGX agent

arXiv:2608.02989v1 Announce Type: cross Abstract: Speculative decoding verifies a tree of draft tokens in one target-model forward pass. For a mixture-of-experts (MoE) target, however, parallel verifi

hardwarearxiv-cs-cl
5 Aug 2026
Research

Activation-Guided Neuron Intervention to Induce Alzheimer's-Related Computational Language Phenotypes in a Large Language Model

DGX agent

arXiv:2608.03067v1 Announce Type: new Abstract: Changes in spontaneous speech provide an early signal of cognitive dysfunction in Alzheimer's disease (AD) that large language models (LLMs) can detect.

researcharxiv-cs-cl
5 Aug 2026
Tutorials

Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

DGX agent

arXiv:2608.03999v1 Announce Type: cross Abstract: Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe,

tutorialsarxiv-cs-cl
5 Aug 2026
← Previous
1…678910…160
Next →