AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
1 Jun 2026

Pairwise Reference Alignment as a Model-Level Ordinal Observable

Model ReleasesDGX agent

arXiv:2605.30758v1 Announce Type: new Abstract: Pairwise preference data is widely used in language-model evaluation and alignment, often for model ranking, reward modeling, or preference optimization

ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs

HardwareDGX agent

arXiv:2602.07721v3 Announce Type: replace-cross Abstract: KV-cache retrieval is essential for long-context LLM inference, yet existing methods struggle with distribution drift and high latency at scal

*-PLUIE: Personalisable metric with Llm Used for Improved Evaluation

SafetyDGX agent

arXiv:2602.15778v2 Announce Type: replace Abstract: Evaluating the quality of automatically generated text often relies on LLM-as-a-judge (LLM-judge) methods. While effective, these approaches are com


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Preference-Aware Rubric Learning for Personalized Evaluation

SafetyDGX agent

arXiv:2605.31545v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve from general-purpose assistants to user-centric agents, personalization has become central to aligning model beha

Probing the Prompt KV Cache: Where It Becomes Dispensable

Model ReleasesDGX agent

arXiv:2605.30574v1 Announce Type: new Abstract: Prior KV cache compression schemes empirically demonstrate that the prompt cache is partially redundant during decoding, dropping or summarising entries

Protocol for evaluating ChatGPT in biomedical association generation and verification using a RAG-enabled, cross-model majority voting workflow

ResearchDGX agent

arXiv:2605.30400v1 Announce Type: new Abstract: We present a protocol to evaluate ChatGPT's ability to generate disease-centric biomedical associations. It outlines how we generate the associations, v

Query-focused and Memory-aware Reranker for Long Context Processing

Model ReleasesDGX agent

arXiv:2602.12192v3 Announce Type: replace Abstract: Built upon the existing analysis of retrieval heads in large language models, we propose an alternative reranking framework that trains models to es

Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses

SafetyDGX agent

arXiv:2504.11972v3 Announce Type: replace Abstract: Extractive QA tasks are commonly evaluated using Exact Match (EM) and F1-score, but these metrics often fail to reflect true model performance. Rece

Refining Word-Based Grammatical Error Annotation for L2 Korean

ResearchDGX agent

arXiv:2605.30545v1 Announce Type: new Abstract: Korean grammatical error correction (K-GEC) presents a structural mismatch between word-based evaluation and the morpheme-level locus of many learner er

Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards

SafetyDGX agent

arXiv:2605.31328v1 Announce Type: new Abstract: Emergent misalignment (EM) is the surprising tendency of language models to become broadly misaligned after fine-tuning on narrowly misaligned examples.

Reliable Multilingual Orthopedic Decision Support from Clinical Narratives: Language-Aware Adaptation and Verification-Guided Deferral

ApplicationsDGX agent

arXiv:2605.31512v1 Announce Type: new Abstract: Multilingual orthopedic decision support remains challenging in low-resource healthcare settings, where clinical narratives contain specialized terminol

Rethinking Sparse Mixture of Experts from a Unified Perspective

ResearchDGX agent

arXiv:2503.22996v3 Announce Type: replace Abstract: Sparse Mixture of Experts (SMoE) models scale the capacity of models while maintaining constant computational overhead. SMoE methods fall into two c

Scaling Multi-Hop Training Data via Graph-Constrained Path Selection

Model ReleasesDGX agent

arXiv:2605.31238v1 Announce Type: new Abstract: Endowing large language models with compositional reasoning over specialized documents requires multi-hop training data at scale, where such data rarely

SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks

ResearchDGX agent

arXiv:2605.31433v1 Announce Type: new Abstract: Self-play can train language models without external supervision. However, existing methods require rule-checkable answers, leaving open-ended tasks dep

Self-Reflective Generation at Test Time

ResearchDGX agent

arXiv:2510.02919v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly solve complex reasoning tasks via long chain-of-thought, but their forward-only autoregressive generation

Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures

SafetyDGX agent

arXiv:2605.30608v1 Announce Type: new Abstract: Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains ch

Semantic Triplet Restoration: A Novel Protocol for Hierarchical Table Understanding in Large Language Models

ResearchDGX agent

arXiv:2605.31550v1 Announce Type: new Abstract: Table question answering requires models to recover semantic relations encoded implicitly by two-dimensional layout, merged cells, and hierarchical head

SERA: Soft-Verified Efficient Repository Agents

Model ReleasesDGX agent

arXiv:2601.20789v3 Announce Type: replace Abstract: Open-weight coding agents should hold a fundamental advantage over closed-source systems because they can specialize to private codebases, encoding

Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents

SafetyDGX agent

arXiv:2605.30723v1 Announce Type: new Abstract: LLM agents increasingly retrieve externally curated skills-procedural instructions retrieved at decision time-to improve performance on long-horizon int

Speculative Decoding Across Languages

ResearchDGX agent

arXiv:2605.30580v1 Announce Type: new Abstract: Speculative decoding has become a crucial component of large language model (LLM) inference, enabling faster generation by drafting multiple tokens and

Speculative Pipeline Decoding: Higher-Accruacy and Zero-Bubble Speculation via Pipeline Parallelism

ResearchDGX agent

arXiv:2605.30852v1 Announce Type: new Abstract: Speculative Decoding (SD) accelerates low-concurrency LLM inference by employing a draft-then-verify paradigm. However, mainstream methods typically rel

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation

SafetyDGX agent

arXiv:2511.11440v3 Announce Type: replace-cross Abstract: Performance gains of Vision Language Models (VLMs) obtained by fine-tuning are generally based on ad hoc data collection and annotation of rea

TaxoBell: Gaussian Box Embeddings for Self-Supervised Taxonomy Expansion

Model ReleasesDGX agent

arXiv:2601.09633v2 Announce Type: replace Abstract: Taxonomies form the backbone of structured knowledge representation across diverse domains, enabling applications such as e-commerce and semantic se

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

Model ReleasesDGX agent

arXiv:2605.30673v1 Announce Type: new Abstract: Classroom videos contain observable teaching practices, but their pedagogical and visual signals are rarely organized in forms suitable for model evalua

The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement

SafetyDGX agent

arXiv:2605.30888v1 Announce Type: new Abstract: Building strong reward models (RMs) for language model alignment is bottlenecked by the cost and difficulty of acquiring diverse and reliable preference

The Latin Substrate: How Language Models Represent and Mediate Script Choice

ResearchDGX agent

arXiv:2605.31363v1 Announce Type: new Abstract: Many languages are written in multiple scripts, requiring large language models (LLMs) to generate equivalent linguistic content in distinct orthographi

The relative strength of hierarchical structure and statistics differs across the measures in naturalistic reading

TutorialsDGX agent

arXiv:2509.23195v2 Announce Type: replace Abstract: The hierarchical syntactic structure and non-hierarchical, statistical, or sequential factors have long been framed as rival theories in accounting

Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining

ApplicationsDGX agent

arXiv:2605.31069v1 Announce Type: cross Abstract: Accurately predicting future events is fundamental to content understanding and decision-making across various domains. While prior research has prima

Towards Efficient LLMs Annealing with Principled Sample Selection

ResearchDGX agent

arXiv:2605.31175v1 Announce Type: new Abstract: The annealing phase is a pivotal convergence stage in LLM pre-training that ultimately determines final model quality. However, effectively selecting tr

TRACE: Discovering Task-Specific Parameter via Adaptation-Aware Probing for Continual Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.31025v1 Announce Type: new Abstract: In real-world deployment, LLMs are often adapted continually across tasks to keep LLMs up-to-date in production, where new fine-tuning should preserve p

Traceable by Design: An LLM Pipeline and Dashboard for EU Regulatory Consultation Analysis

SafetyDGX agent

arXiv:2605.30995v1 Announce Type: cross Abstract: Public consultations generate large volumes of data in the form of stakeholder submissions that are practically unfeasible to analyse manually. We pre

Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing

TutorialsDGX agent

arXiv:2605.31367v1 Announce Type: cross Abstract: Token mixing layers play a key role in how language models can learn and generate long-range dependencies. Their efficiency relies on the necessary tr

Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation Workflows

Model ReleasesDGX agent

arXiv:2605.31452v1 Announce Type: new Abstract: Building on our previous work, this paper develops practical, low-barrier methods for freelance translators and smaller language service providers to ev

TransLPRNet: Lite Vision-Language Network for Single/Dual-line Chinese License Plate Recognition

ResearchDGX agent

arXiv:2507.17335v2 Announce Type: replace-cross Abstract: License plate recognition in open environments is widely applicable across various domains; however, the diversity of license plate types and

Triaging Threats to Specialized Guardrails

Model ReleasesDGX agent

arXiv:2605.30693v1 Announce Type: cross Abstract: Building robust safety guardrails is essential for deploying Large Language Models across diverse real-world applications. However, this goal remains

TSM-Bench: Detecting LLM-Generated Text in Real-World Wikipedia Editing Practices

Model ReleasesDGX agent

arXiv:2605.31113v1 Announce Type: new Abstract: Automatically detecting machine-generated text (MGT) is critical to maintaining the knowledge integrity of user-generated content (UGC) platforms such a

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception

SafetyDGX agent

arXiv:2605.31521v1 Announce Type: new Abstract: Semantic speech tokenizers have become a widely used interface for Audio-LLMs, owing to their compact single-codebook design and strong linguistic align

UniDial-EvalKit: A Unified Toolkit for Evaluating Multi-Faceted Conversational Abilities

Model ReleasesDGX agent

arXiv:2603.23160v2 Announce Type: replace Abstract: Benchmarking large language models (LLMs) and agents in multi-turn interactive scenarios is essential for understanding their practical capabilities

Unlocking Fine-Grained Translation Quality Estimation in LRMs through Synergistically Evolving Implicit and Explicit Reasoning

ResearchDGX agent

arXiv:2605.31378v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) still struggle with fine-grained translation quality estimation (QE), even with long reasoning chains. We argue that LRMs

Weights to Code: Extracting Interpretable Algorithms from the Discrete Transformer

ResearchDGX agent

arXiv:2601.05770v3 Announce Type: replace-cross Abstract: Algorithm extraction aims to synthesize executable programs directly from models trained on algorithmic tasks, enabling de novo recovery of ex

What Am I Missing? Question-Answering as Hidden State Probing

SafetyDGX agent

arXiv:2605.31561v1 Announce Type: new Abstract: Test-time reasoning has become a significant field of study since the introduction of chain-of-thought reasoning in large language models (LLMs). Howeve

When English Rewrites Local Knowledge: Global Narrative Dominance in Large Language Models

Local AiDGX agent

arXiv:2605.30481v1 Announce Type: new Abstract: Large language models (LLMs) are widely used as cross-lingual knowledge interfaces. However, culturally grounded questions often reflect globally domina

Wind Turbine Maintenance Log Labelling Framework: LLM-Driven Data Correction and Enrichment via Semantic Extraction of Reliability Intelligence

ResearchDGX agent

arXiv:2605.31281v1 Announce Type: new Abstract: As wind turbine fleets age, data-driven reliability engineering is essential to optimise their operation and maintenance for service life extension and

Your Multimodal Speech Model Says I Have a Face for Radio

Model ReleasesDGX agent

arXiv:2605.30472v1 Announce Type: new Abstract: As large neural models have become better at language tasks, researchers are increasingly building multi- and omnimodal models that handle more modaliti

29 May 2026

A Dual-Path Architecture for Scaling Compute and Capacity in LLMs

Model ReleasesDGX agent

arXiv:2605.30202v1 Announce Type: new Abstract: Looped transformers apply a shared block multiple times and have emerged as a parameter-efficient route to scaling compute in language models. However,

A Modular Architecture for Typologically Controlled Lexicon Generation

SafetyDGX agent

arXiv:2605.28824v1 Announce Type: new Abstract: Constructing artificial lexicons that are pronounceable, typologically plausible, and semantically structured remains an open challenge in computational

A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities

SafetyDGX agent

arXiv:2605.29340v1 Announce Type: new Abstract: In this paper, we discuss question-answer dataset for LLM safety evaluation, with a focus on illegal activities. Specifically, on the basis of manual an

Accommodation Goes Both Ways: Studying Linguistic Convergence Between Humans and Language Models

ApplicationsDGX agent

arXiv:2605.29278v1 Announce Type: new Abstract: As LLMs become increasingly integrated into daily life, understanding how their presence will shape human linguistic behavior is an open question. We pr

ActTraitBench: Quantifying the Knowledge-Decision Gap in Large Language Models via Human-Grounded Behavioral Validation

SafetyDGX agent

arXiv:2605.29791v1 Announce Type: new Abstract: While Large Language Models (LLMs) can convincingly simulate personas in explicit self-reports, they often deviate in implicit behavioral decisions, rev

Adapting Multilingual Embedding Models to Turkish via Cross-Lingual Tokenizer Surgery and Offline Distillation

Model ReleasesDGX agent

arXiv:2605.29992v1 Announce Type: new Abstract: Sentence embeddings are a foundational component for semantic search, clustering, classification, and retrieval-augmented generation. This paper present

Adaptive Targeted Dynamic Chunking for Tokenization-Free Hierarchical Model

ResearchDGX agent

arXiv:2605.30080v1 Announce Type: new Abstract: Tokenization-free hierarchical models are emerging as a promising alternative to traditional Large Language Models (LLMs), addressing inherent preproces

AfriScience-MT: Towards Decolonizing Science in Africa through Text Translation

Model ReleasesDGX agent

arXiv:2605.29741v1 Announce Type: new Abstract: The dominance of colonial languages in African education and scientific communication limits how hundreds of millions of speakers of African languages a

Analyzing Persona Effects in Generated Explanations from Multimodal LLM Agents in Urban Perception

ResearchDGX agent

arXiv:2605.29064v1 Announce Type: new Abstract: We study how persona prompting shapes language generated by multimodal large language models in an urban perception setting. Using 59,808 annotations fr

Attention Asymmetry in AI Layoff Discourse on X: A Computational Analysis of Capital vs Labour Amplification

ResearchDGX agent

arXiv:2605.29367v1 Announce Type: new Abstract: When workers lose jobs to AI-driven restructuring, two very different conversations happen on X (formerly Twitter) at the same time. Tech executives and

'Be My Cheese?': Cultural Nuance Benchmarking for Machine Translation in Multilingual LLMs

Model ReleasesDGX agent

arXiv:2602.04729v2 Announce Type: replace Abstract: We present a large-scale human evaluation benchmark for assessing cultural localisation in machine translation produced by state-of-the-art multilin

Beyond English and Evasion: A Human-Annotated Multi-Domain Benchmark for High-Stakes LLM Safety Evaluation in Chinese

Model ReleasesDGX agent

arXiv:2605.29667v1 Announce Type: new Abstract: When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break

Beyond Transcripts: A Renewed Perspective on Audio Chaptering

ResearchDGX agent

arXiv:2602.08979v2 Announce Type: replace-cross Abstract: Audio chaptering, the task of segmenting long-form audio into coherent sections, is increasingly important for navigating podcasts, lectures,

Bosses, Kings, and the Commons: Cooperation Under Power Asymmetry in LLM Societies

AgentsDGX agent

arXiv:2605.29062v1 Announce Type: new Abstract: Communities can sustainably manage shared resources (commons) through self-governance and cooperative norms, a central finding of Ostrom's theory of sel

BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base

Model ReleasesDGX agent

arXiv:2605.29379v1 Announce Type: new Abstract: We present BrahmicTokenizer-131K, a 131,072-vocabulary byte-level BPE tokenizer that closes the Brahmic compression gap at the 131K-vocabulary class whi

Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations

SafetyDGX agent

arXiv:2601.08064v2 Announce Type: replace Abstract: Confidence estimation (CE) indicates how reliable the answers of large language models are and impacts user trust and decision-making. Existing eval

← Previous
1…5253545556…129
Next →