AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
6 Aug 2026

State2State: Environment-Derived Mid-Training for LLM Agents

AgentsDGX agent

arXiv:2608.04934v1 Announce Type: new Abstract: Training LLM agents commonly relies on supervised fine-tuning from expert trajectories or online reinforcement learning over human-specified tasks with

Strengthening Target-Language Features: SAE-Based Steering for Multilingual Inference

Model ReleasesDGX agent

arXiv:2608.04904v1 Announce Type: new Abstract: Multilingual large language models exhibit substantial performance differences across languages, while existing adaptation methods often require paramet

STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation

Model ReleasesDGX agent

arXiv:2608.04567v1 Announce Type: new Abstract: Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language p


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Test, then Route: How Language Models Execute In-Context Conditional Rules Across Models and Languages

Model ReleasesDGX agent

arXiv:2608.04183v1 Announce Type: new Abstract: When a language model follows an in-context conditional rule such as 'if P(x) then A else B,' does it assemble a runtime circuit with one module that te

The Calibration Floor: Format Repair Can Masquerade as Self-Correction at Small-to-Mid Scale

Model ReleasesDGX agent

arXiv:2608.04355v1 Announce Type: new Abstract: Accuracy changes after language-model self-revision are usually interpreted as changes in reasoning. We show this can fail at the answer-extraction boun

The Evaluator Is Part of the Experiment: Measuring Open-Ended LLM Conformity

Model ReleasesDGX agent

arXiv:2608.04463v1 Announce Type: new Abstract: Prior work on LLM conformity largely measures discrete answer flips under verifiable labels. Open-ended revisions require a different measurement strate

The Fairness Collapse Phenomenon: Bias Amplification in Language Models Trained on Synthetic Data

SafetyDGX agent

arXiv:2608.04268v1 Announce Type: new Abstract: Generative models trained on artificially generated data have been shown to exhibit model collapse, resulting in significant performance degradation. As

The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

ResearchDGX agent

arXiv:2608.04570v1 Announce Type: new Abstract: Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inferenc

TopoChunker: Topology-Aware Agentic Document Chunking Framework

AgentsDGX agent

arXiv:2603.18409v2 Announce Type: replace Abstract: Current document chunking methods for Retrieval-Augmented Generation (RAG) typically linearize text. This forced linearization strips away intrinsic

Toward Federated Large Language Models in Medicine: A Parameter-Efficient Framework for Privacy-Preserving, Multi-Institutional Adaptation

Model ReleasesDGX agent

arXiv:2601.22124v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly adapted for medical applications, but most are trained using data from a single institution because pr

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

Model ReleasesDGX agent

arXiv:2608.05139v1 Announce Type: new Abstract: Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivat

Towards End-to-End Multilingual Metaphor Processing: Integrating Detection, Translation, and Evaluation

ResearchDGX agent

arXiv:2608.04260v1 Announce Type: new Abstract: Metaphorical language remains a major challenge for multilingual natural language processing because successful interpretation and translation require r

Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs

Local AiDGX agent

arXiv:2608.04759v1 Announce Type: cross Abstract: Although Multimodal Large Language Models (MLLMs) have made substantial progress, their spatial reasoning may still produce intermediate judgments inc

Transfer Learning for Named Entity Recognition of Classical Latin through LLM Prompting

Model ReleasesDGX agent

arXiv:2608.04015v1 Announce Type: new Abstract: With the increase in digitized resources of Classical Latin texts and modern breakthroughs of Large Language Models (LLMs), I contribute to ancient lang

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

Model ReleasesDGX agent

arXiv:2607.19322v2 Announce Type: replace Abstract: Rubric-based evaluation of open-ended generation faces a fundamental tension between expressiveness and reliability. Authoring a faithful rubric req

UG-UMRE: Uncertainty-Guided Modality Augmentation and Distributional Calibration for Unified Multimodal Relation Extraction

Model ReleasesDGX agent

arXiv:2608.04949v1 Announce Type: cross Abstract: Unified Multimodal Relation Extraction (UMRE) aims to identify intra-modal and cross-modal relations between textual entities and visual objects. Howe

Unforgettable Generalization in Language Models

TutorialsDGX agent

arXiv:2409.02228v2 Announce Type: replace-cross Abstract: When language models (LMs) are trained to forget (or 'unlearn'') a skill, how precisely does their behavior change? We study the behavior of t

When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents

SafetyDGX agent

arXiv:2608.04574v1 Announce Type: new Abstract: Memory-augmented VLM agents act on persistent spatial knowledge, yet that knowledge silently goes stale as the environment changes. We ask what happens

When Modalities Remember: Continual Learning for Multimodal Knowledge Graphs

SafetyDGX agent

arXiv:2604.02778v2 Announce Type: replace Abstract: Real-world multimodal knowledge graphs (MMKGs) are dynamic, with new entities, relations, and multimodal knowledge emerging over time. Existing cont

When More Becomes Less: Position-Dependent Repetition Effects in Language Models

ResearchDGX agent

arXiv:2608.04021v1 Announce Type: new Abstract: Cloze-style probes that vary how often a target token appears implicitly assume that more copies of a target affect prediction the same way regardless o

5 Aug 2026

A machine-readable catalogue of the Tsiolkovsky papers (fond 555, Archive of the Russian Academy of Sciences), and a way to measure how well its handwriting can be read

ResearchDGX agent

arXiv:2608.03617v1 Announce Type: new Abstract: The personal archive of Konstantin Tsiolkovsky (1857-1935) is held as fond 555 of the Archive of the Russian Academy of Sciences. The archive scanned th

AcceptMoE: Commitment-Weighted Self-Sizing Verifier Expert Sets for Efficient MoE Speculative Decoding

HardwareDGX agent

arXiv:2608.02989v1 Announce Type: cross Abstract: Speculative decoding verifies a tree of draft tokens in one target-model forward pass. For a mixture-of-experts (MoE) target, however, parallel verifi

Activation-Guided Neuron Intervention to Induce Alzheimer's-Related Computational Language Phenotypes in a Large Language Model

ResearchDGX agent

arXiv:2608.03067v1 Announce Type: new Abstract: Changes in spontaneous speech provide an early signal of cognitive dysfunction in Alzheimer's disease (AD) that large language models (LLMs) can detect.

Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

TutorialsDGX agent

arXiv:2608.03999v1 Announce Type: cross Abstract: Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe,

Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech

SafetyDGX agent

arXiv:2608.02941v1 Announce Type: new Abstract: We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-C

An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures

AgentsDGX agent

arXiv:2608.03735v1 Announce Type: cross Abstract: Multilingual multi-agent systems exhibit substantial degradation beyond English, yet prior work rarely identifies how task-critical information is los

ANCHOR-RE: An Agentic Neuro-Symbolic Framework for Grounded Biomedical Relation Extraction

Model ReleasesDGX agent

arXiv:2608.03154v1 Announce Type: new Abstract: Biomedical relation extraction (BioRE) extracts structured knowledge from biomedical literature for applications such as knowledge base construction and

AnchorKV: Anchor-Residual KV Cache Compression

ResearchDGX agent

arXiv:2608.02901v1 Announce Type: cross Abstract: The key-value (KV) cache is the primary memory bottleneck in long-context LLM inference. Existing approaches attack it from opposite ends: eviction me

ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts

Model ReleasesDGX agent

arXiv:2608.03898v1 Announce Type: new Abstract: The automatic structural analysis of legal texts is a cornerstone of legal technology, yet the extraction of their logical components remains a signific

ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads

ResearchDGX agent

arXiv:2608.02703v1 Announce Type: new Abstract: Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but practical backends often retain the fin

ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2608.03358v1 Announce Type: new Abstract: Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emot

ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference

Model ReleasesDGX agent

arXiv:2608.02947v1 Announce Type: cross Abstract: The attention score with rotary position embeddings (RoPE) decomposes exactly into a sum over its 2D-rotation frequency pairs, and each pair's wavelen

Attention is Case-Sensitive

ResearchDGX agent

arXiv:2608.03711v1 Announce Type: cross Abstract: In human visual perception, uppercase lettering serves as a natural salience cue that captures attention within lowercase text. In this paper, we pres

BanglaWild: An In-the-Wild Bengali Scene Text Recognition Benchmark for OCR and Vision-Language Models

Model ReleasesDGX agent

arXiv:2608.03884v1 Announce Type: cross Abstract: In-the-wild Bengali scene text recognition is largely unmeasured: existing resources target handwritten documents or constrained sign-board parsing, r

BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems

Model ReleasesDGX agent

arXiv:2608.02612v1 Announce Type: new Abstract: Formulating an optimization problem strongly affects the quality of the final solution, yet good formulations usually require substantial expertise. Rec

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks

ApplicationsDGX agent

arXiv:2608.03340v1 Announce Type: new Abstract: Predicting LLM's capabilities on real-world tasks is essential, yet the extent to which performance on commonsense benchmarks predicts downstream perfor

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

ResearchDGX agent

arXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how m

Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension

Model ReleasesDGX agent

arXiv:2608.03494v1 Announce Type: new Abstract: Vocabulary extension is an efficient way to adapt pretrained large language models (LLMs) to new languages, but the initialization of newly added token

Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

Model ReleasesDGX agent

arXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive

BOW: Training Language Models to Reason Over Plausible Next Words

SafetyDGX agent

arXiv:2506.13502v3 Announce Type: replace Abstract: Next-word prediction (NWP) trains language models against a single observed continuation, even though many contexts admit multiple plausible next wo

Calibrating Semantic Uncertainty from Observable Language-Model Probabilities

ResearchDGX agent

arXiv:2607.17447v2 Announce Type: replace-cross Abstract: As generative artificial intelligence enters scientific and professional work, its uncertainty must be defined on the states that matter for i

Character Iconicity vs. Arbitrariness: An Arabic NLP Perspective

ResearchDGX agent

arXiv:2608.02935v1 Announce Type: new Abstract: Arabic script uses 28 letters, many of which share a common base shape (rasm) and are distinguished only by dot placement. Because early Arabic manuscri

CIGTSurv: Clinical Information Guided Tri-modal Survival Prediction with Local Prototype Association and Global Feature Alignment

Local AiDGX agent

arXiv:2608.03247v1 Announce Type: cross Abstract: Multimodal learning has significantly advanced survival prediction by integrating pathology images with genomic data. However, clinical information, d

ConlangBench: Exploring Language Knowledge and Learning in LLMs through Diverse Constructed Languages

Model ReleasesDGX agent

arXiv:2608.03505v1 Announce Type: new Abstract: Constructed languages (conlangs) are intentionally created human languages with a rich tradition of linguistic creativity. Despite their potential for s

Consensus Measures for Unstructured Biomedical Text Annotations

ResearchDGX agent

arXiv:2608.03529v1 Announce Type: new Abstract: Biomedical literature is increasingly mined for knowledge beyond the questions it was written to answer. Because the target concepts are not known in ad

Convex-Hull-Neighborhood Smooth Dual Generalization: Controlling Local Correction Propagation in Offline RL

ResearchDGX agent

arXiv:2608.03108v1 Announce Type: cross Abstract: Offline reinforcement learning (offline RL) can benefit from nearby out-of-distribution (OOD) actions, but estimation errors at these actions may be a

Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation

ApplicationsDGX agent

arXiv:2608.02694v1 Announce Type: new Abstract: Long-horizon video editing agents receive final-product feedback only after many interdependent decisions. Yet editing quality is subjective, admits mul

Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili

Model ReleasesDGX agent

arXiv:2608.03532v1 Announce Type: new Abstract: Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric

Detecting Hallucinations and Recovering Verified Answers in Arabic Islamic Question Answering

Model ReleasesDGX agent

arXiv:2608.03720v1 Announce Type: new Abstract: Large language models can generate fluent responses to Islamic questions while introducing factual errors that are difficult to identify. This paper pre

Disentangling Language Modeling and Boundaries

SafetyDGX agent

arXiv:2608.03599v1 Announce Type: new Abstract: Byte-level language models are usually argued for on the grounds of robustness, multilingual fairness, and character-level skills. We point to a differe

Disentangling MLP Neuron Weights in Vocabulary Space

Model ReleasesDGX agent

arXiv:2604.06005v2 Announce Type: replace Abstract: Interpreting the information encoded in language model weights remains a fundamental challenge in mechanistic interpretability. In this work, we int

Don't Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference

ResearchDGX agent

arXiv:2608.03388v1 Announce Type: new Abstract: Abductive reasoning requires forming hypotheses that explain observed evidence and revising them as new evidence becomes available. While large language

Don't Walk the Line: Boundary Guidance for Filtered Generation

Model ReleasesDGX agent

arXiv:2510.11834v3 Announce Type: replace-cross Abstract: Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. A common strategy is to fine-tun

DP-MemView: A Memory Interface for Attribute-Level Transcript Privacy in Long-Term LLM Agents

Model ReleasesDGX agent

arXiv:2608.03130v1 Announce Type: cross Abstract: Long-term memory enables persistent personalization in LLM agents, but repeated memory-conditioned responses can cumulatively reveal protected attribu

DS@GT-ARC at eRisk 2026 Task 3: Sparse, Semantic, and LLM Reranking for ADHD Symptom Sentences

Model ReleasesDGX agent

arXiv:2608.03883v1 Announce Type: new Abstract: This paper describes our submissions to eRisk 2026 Task 3, ADHD Symptom Sentence Ranking. The task requires systems to rank candidate Reddit sentences a

DUD: Decoupled Update Dynamics for Reliable Uncertainty Quantification in Large Language Models

ResearchDGX agent

arXiv:2608.03411v1 Announce Type: new Abstract: Accurate Uncertainty Quantification (UQ) is critical for reliable deployment of Large Language Models (LLMs), yet traditional probability-based metrics

Dynamically Allocating Evaluation Effort for Model Ranking

Model ReleasesDGX agent

arXiv:2608.03437v1 Announce Type: new Abstract: While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing m

Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study

Model ReleasesDGX agent

arXiv:2608.03480v1 Announce Type: new Abstract: The adoption of large pre-trained multilingual models for neural machine translation (MNMT) faces a major challenge: excessive memory and computational

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks

ResearchDGX agent

arXiv:2608.02966v1 Announce Type: new Abstract: Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scorin

FLARE: Few-shot Learning-based Adaptive Reflective Engine

Model ReleasesDGX agent

arXiv:2608.02919v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in complex, compound AI systems where performance hinges on the quality of prompts. Recent state-

← Previous
1…56789…128
Next →