AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
5 Aug 2026

From SQL Errors to Concept Gaps: An AI-Powered Knowledge Graph Analytics Platform for Personalized Feedback

TutorialsDGX agent

arXiv:2608.03118v1 Announce Type: new Abstract: This innovative practice full paper describes an AI-powered knowledge graph platform that connects SQL errors to conceptual gaps in undergraduate and gr

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models

Local AiDGX agent

arXiv:2608.03083v1 Announce Type: cross Abstract: Video large language models (VideoLLMs) achieve strong video understanding performance, but their inference remains expensive due to the large number

HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification

Research

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2608.03966v1 Announce Type: new Abstract: Large language models can generate fluent Arabic answers while introducing factual errors that are difficult to identify and verify. Existing Arabic hal

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning

SafetyDGX agent

arXiv:2608.03545v1 Announce Type: new Abstract: Test-time reinforcement learning (TTRL) improves the reasoning capabilities of large language models without labeled data by updating the policy with ps

HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection

Model ReleasesDGX agent

arXiv:2509.23690v2 Announce Type: replace-cross Abstract: Safety hazards in the home are a leading cause of preventable domestic injuries, motivating an automated inspector that actively explores a ho

HomoEnsNER: Does Language Alignment Outperform Architectural Complexity in Gujarati Named Entity Recognition?

SafetyDGX agent

arXiv:2608.03105v1 Announce Type: new Abstract: Named Entity Recognition (NER) for Gujarati remains underexplored, hindered by the absence of capitalization cues, rich morphology, lexical ambiguity, a

HUKUKBERT: Domain-Specific Language Model for Turkish Law

Model ReleasesDGX agent

arXiv:2604.04790v2 Announce Type: replace Abstract: Natural language processing (NLP) advances have powered a generation of LegalTech systems, but Turkish law remains under-served by domain-specific d

ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization

SafetyDGX agent

arXiv:2608.03210v1 Announce Type: new Abstract: Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift

JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation

Model ReleasesDGX agent

arXiv:2608.02620v1 Announce Type: new Abstract: LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their o

LACE: Large Language Model Aided Multi-Agent Framework for Agile RISC-V Instruction Extension

AgentsDGX agent

arXiv:2608.02915v1 Announce Type: cross Abstract: Domain-specific Instruction Set Architecture eXtensions (ISAX) are widely adopted in the RISC-V ecosystem to accelerate emerging workloads, but implem

Language Models Encode the Contextual Truth of Propositions

ResearchDGX agent

arXiv:2608.03035v1 Announce Type: new Abstract: Prior work has shown that LLMs encode the truth of factual propositions along linear directions in activation space. It's unclear how these representati

Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR

SafetyDGX agent

arXiv:2608.03610v1 Announce Type: new Abstract: Modern LLM-based ASR systems have established multilingual capability as a standard feature, leveraging large-scale multilingual corpora and LLMs' cross

Looking under the Wrong Lamppost: On the Limitations of Automated Translation Quality Estimation

Model ReleasesDGX agent

arXiv:2608.03577v1 Announce Type: new Abstract: Automation of Translation Quality Estimation (QE) has emerged as a widely discussed approach to managing translation quality at scale, and a growing num

LoopMTP: A looped transformer guided by latent multi-token prediction

Model ReleasesDGX agent

arXiv:2608.03624v1 Announce Type: new Abstract: Looped transformers have emerged as a parameter-efficient alternative to scaling depth for strong reasoning. By reusing one stack of layers across T ite

M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models

Model ReleasesDGX agent

arXiv:2608.03803v1 Announce Type: new Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a languag

Mapping the City Through the Lens of Language Models

ResearchDGX agent

arXiv:2608.02971v1 Announce Type: new Abstract: Language models often complete an underspecified reference to a city with unstated assumptions about urban size, form, infrastructure, environment, and

MIDI-LLM: Improving Text-to-MIDI Music Generation via Adapting Large Language Models

Model ReleasesDGX agent

arXiv:2511.03942v2 Announce Type: replace-cross Abstract: We present MIDI-LLM, a recipe that improves multitrack text-to-MIDI generation via adapting Large Language Models (LLMs). MIDI-LLM expands an

MoEGen: Mixture-of-Experts for Instance-Adaptive LoRA Generation

Model ReleasesDGX agent

arXiv:2608.03275v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of large language models, but existing MoE-based PEFT methods typically improve capa

On the Diversity of Analogy Making in Large Language Models

ResearchDGX agent

arXiv:2608.03233v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable potential for analogy making, a core cognitive capability that drives novelty and creativity.

On the Non-Specificity of Statistical Measures Used in Script Decipherment

ResearchDGX agent

arXiv:2608.02999v1 Announce Type: new Abstract: Statistical regularities are routinely offered as evidence that undeciphered sign systems encode language; the Indus script debate is the canonical exam

OPTD: On-Policy Transition Distillation with Consistency-Guided Adaptive Compression for Few-Step Diffusion Language Models

SafetyDGX agent

arXiv:2608.02942v1 Announce Type: new Abstract: Diffusion language models (dLLMs) can predict many tokens in parallel, but accurate generation still requires many iterative denoising steps. Few-step d

PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation

ResearchDGX agent

arXiv:2608.03077v1 Announce Type: new Abstract: Multi-domain machine translation (MDMT) requires more than fluent generation: it demands domain-sensitive translation decisions such as domain disambigu

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs

Model ReleasesDGX agent

arXiv:2608.04010v1 Announce Type: cross Abstract: Existing scaling strategies for Multimodal Large Language Models (MLLMs) typically expand either model parameters or sequential inference computation,

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents

Model ReleasesDGX agent

arXiv:2608.04003v1 Announce Type: new Abstract: Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for s

PDD-RRG: Posterior Diagnostic Decision for Study-level Radiology Report Generation

ResearchDGX agent

arXiv:2608.03055v1 Announce Type: cross Abstract: Automatic radiology report generation (RRG) aims to simulate the workflow of radiologists, assisting them in clinical diagnosis. However, existing met

Pingala: Prosody-Aware Decoding for Sanskrit Poetry Generation

Model ReleasesDGX agent

arXiv:2603.24413v2 Announce Type: replace Abstract: Poetry generation in Sanskrit typically requires the verse to be semantically coherent and adhere to strict prosodic rules. In Sanskrit prosody, eve

Predicting Deep Neural Network Training Outcomes from Early Training Telemetry

ResearchDGX agent

arXiv:2608.03709v1 Announce Type: new Abstract: Large hyperparameter sweeps for deep neural networks spend substantial compute on configurations that are effectively doomed from the first few epochs.

Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment nicode{x2013} Is English Enough?

SafetyDGX agent

arXiv:2608.03446v1 Announce Type: new Abstract: Multilingual large language models (LLMs) have been shown to perform better on non-English classification tasks when the representations of the given la

Probing Character-level Transformers for the Spanish L-shaped Morphome

Local AiDGX agent

arXiv:2608.03452v1 Announce Type: new Abstract: When a transformer learns an irregular morphological pattern, what has it learned? Our test case is the Spanish L-shaped morphome, a complex irregular p

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning

SafetyDGX agent

arXiv:2608.02831v1 Announce Type: cross Abstract: Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning

Relational Priors as Convergence Pressure in LLM-Based Multi-Agent Systems

SafetyDGX agent

arXiv:2608.03239v1 Announce Type: new Abstract: Large language model-based multi-agent systems (LLM-MAS) are designed through roles, debate protocols, and aggregation rules. These choices create impli

Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining

SafetyDGX agent

arXiv:2608.03089v1 Announce Type: new Abstract: Large-scale pretraining corpora contain substantial duplicate content. Although document-level deduplication is widely used, removing subdocument-level

Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling

Model ReleasesDGX agent

arXiv:2608.03842v1 Announce Type: new Abstract: When a language model fails on surface-perturbed input (typos, OCR noise, homophones), 'which layer is responsible' has three natural operationalization

SeqLLM: Augmenting LLMs with Behavioral-Sequence Modeling for High-Stakes Decisions at WeChat Pay

Model ReleasesDGX agent

arXiv:2608.03063v1 Announce Type: new Abstract: Merchant risk control at large payment platforms screens tens of millions of merchants daily, where false positives harm legitimate merchants and false

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs

Model ReleasesDGX agent

arXiv:2608.03573v1 Announce Type: new Abstract: Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large langu

SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA

Model ReleasesDGX agent

arXiv:2509.25459v2 Announce Type: replace Abstract: Large Language Models (LLMs) show promise in generating long-form scientific explanations that synthesize evidence and connect multiple factors. How

SocietyBench: Forecasting Counterfactual Social-World Evolution

Model ReleasesDGX agent

arXiv:2608.04009v1 Announce Type: new Abstract: Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a b

Sparse Weight Decomposition for Efficient Circuit Extraction

TutorialsDGX agent

arXiv:2608.03913v1 Announce Type: cross Abstract: Dense pretrained transformers do not naturally expose interpretable units for circuit extraction. Existing approaches obtain such units by learning au

States Hidden in Hidden States: Implicit Discrete State Representations Emerge in LLMs' Hidden States

ResearchDGX agent

arXiv:2407.11421v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit emergent abilities that may reveal aspects of their internal mechanisms. We study one such capability: directly

string2string Studio: An Interactive, In-Browser Platform for String-to-String Algorithms

SafetyDGX agent

arXiv:2608.03984v1 Announce Type: new Abstract: We present string2string Studio, an interactive in-browser platform for string-to-string analysis across natural language processing, computational biol

Stuck on 'A': Diagnosing and Repairing Interface Injury in Attention-to-KDA Linearization of a 0.6B Language Model

Model ReleasesDGX agent

arXiv:2608.02689v1 Announce Type: new Abstract: We convert 21 of 28 full-attention layers of Qwen3-0.6B-Base into KDA (Kimi Delta Attention) linear-attention layers on a single consumer-grade GPU budg

Suffix-Constrained Greedy Search Algorithms for Causal Language Models

ResearchDGX agent

arXiv:2603.01243v2 Announce Type: replace Abstract: Large language models (LLMs) are powerful tools that have found applications beyond human-machine interfaces and chatbots. Beside free-form generati

TabletCraft: Bridging a 4,000-Year Cultural Gap with Bidirectional Akkadian NMT and Cuneiform Rendering

ResearchDGX agent

arXiv:2608.02609v1 Announce Type: new Abstract: Half a million cuneiform clay tablets survive in museums worldwide, yet modern users can neither read nor write in the world's oldest writing system, le

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores

ResearchDGX agent

arXiv:2608.02985v1 Announce Type: cross Abstract: The standard check for contamination in LLM backtests is simple: compare scores before and after the training cutoff. We show this check is uninformat

The Eloquence team submission for task 1 of MLC-SLM challenge

ApplicationsDGX agent

arXiv:2507.19308v2 Announce Type: replace-cross Abstract: In this paper, we present our studies and experiments carried out for the task 1 of the Challenge and Workshop on Multilingual Conversational

VetScore: Risk-Weighted Fact Verification for Veterinary Long-Form QA with Citations

ResearchDGX agent

arXiv:2608.03675v1 Announce Type: new Abstract: Citation excerpts can be used to increase the reliability of generated outputs and their faithfulness to cited sources, which is especially important in

VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP

Model ReleasesDGX agent

arXiv:2608.03095v1 Announce Type: new Abstract: We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurati

What Language Does and What the Evidence Supports: A Functional Role Taxonomy and Evidence Audit of Language Grounding in Embodied Agents

ResearchDGX agent

arXiv:2608.03099v1 Announce Type: new Abstract: Foundation models place language throughout embodied agents, but its presence does not show what it contributes or how well that contribution is grounde

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

Model ReleasesDGX agent

arXiv:2608.03700v1 Announce Type: cross Abstract: Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personaliz

When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings

Model ReleasesDGX agent

arXiv:2608.03994v1 Announce Type: new Abstract: We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows floating-point precision, which zeroes

Word Recovery in Large Language Models Enables Character-Level Tokenization Robustness

ResearchDGX agent

arXiv:2603.10771v2 Announce Type: replace Abstract: Large language models (LLMs) trained with canonical tokenization exhibit surprising robustness to non-canonical inputs such as character-level token

WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament

Model ReleasesDGX agent

arXiv:2608.04008v1 Announce Type: new Abstract: Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewher

4 Aug 2026

A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models

ResearchDGX agent

arXiv:2509.22536v5 Announce Type: replace Abstract: The immense computational cost of training Large Language Models (LLMs) presents a major barrier to innovation. While FP8 training offers a promisin

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard)

SafetyDGX agent

arXiv:2608.00180v1 Announce Type: new Abstract: Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard. Training a safety guard with RL means optimizing two

A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense

AgentsDGX agent

arXiv:2608.00583v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring is meant to catch the reward hacks that look clean in the actions and betray themselves only in the reasoning. We sh

A Few Neurons Reveal When LLMs Misuse Tools: Sparse Detection and Selective Steering for Reliable Tool Use

Model ReleasesDGX agent

arXiv:2608.00218v1 Announce Type: new Abstract: Agentic LLMs exhibit three consequential tool-use failures: invalid arguments (validity), unnecessary calls (over-calling), and omitted calls when tools

A Fortran General-Purpose Transpiler: Proof of Concept

HardwareDGX agent

arXiv:2608.00130v1 Announce Type: cross Abstract: Fortran has been the cornerstone of high-performance computing for decades and remains unmatched in many domains. Yet the language faces an expertise

A Heuristic Perspective on Debiasing Language Models

SafetyDGX agent

arXiv:2608.00622v1 Announce Type: new Abstract: Language models (LMs) often acquire various biases during pre-training and may express them in interactions, potentially causing social harm. Existing m

A Large-Scale Multi-Dimensional Empirical Study of LLMs for Conversation Summarization

Model ReleasesDGX agent

arXiv:2606.15974v2 Announce Type: replace Abstract: Despite the significant advancement of LLMs in conversation summarization, their evaluation remains limited by insufficient scenarios, input lengths

A Triple-Robustness Analysis of Retrieval-Augmented Generation for Multi-Hop Requirements Traceability

Model ReleasesDGX agent

arXiv:2608.00705v1 Announce Type: cross Abstract: Reported verdicts on GraphRAG versus vector RAG disagree, and the evidence is typically tied to a single corpus, embedder, and judge -- and, we show,

← Previous
1…678910…128
Next →