AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

MADS: Model-Aware Diverse Core Set Selection for Instruction Tuning

DGX agent

arXiv:2605.30857v1 Announce Type: new Abstract: Instruction fine-tuning is employed to enhance the instruction-following ability of large language models (LLMs). As the amount of instruction fine-tuni

model-releasesarxiv-cs-cl
1 Jun 2026
Local Ai
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Measuring, Localizing, and Ablating Alignment Signatures in LLMs

DGX agent

arXiv:2605.30526v1 Announce Type: cross Abstract: Aligned language models often exhibit a recognizable AI-like style, yet its connection to post-training and internal representations remains poorly un

local-aiarxiv-cs-cl
1 Jun 2026
Model Releases

Mellum2 Technical Report

DGX agent

arXiv:2605.31268v1 Announce Type: new Abstract: We present Mellum 2, an open-weight 12B-parameter Mixture-of-Experts (MoE) language model with 2.5B active parameters per token. Mellum 2 is a general-p

model-releasesarxiv-cs-cl
1 Jun 2026
Local Ai

Memory-Efficient Structured Backpropagation for On-Device LLM Fine-Tuning

DGX agent

arXiv:2602.13069v2 Announce Type: replace-cross Abstract: On-device fine-tuning enables privacy-preserving personalization of large language models, but mobile devices impose severe memory constraints

local-aiarxiv-cs-cl
1 Jun 2026
Model Releases

MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft

DGX agent

arXiv:2605.30931v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to susta

model-releasesarxiv-cs-cl
1 Jun 2026
Research

MoG: Mixture of Experts for Graph-based Retrieval-Augmented Generation

DGX agent

arXiv:2605.31010v1 Announce Type: new Abstract: Retrieval-augmented generation is intensively studied to ground large language models on external evidence. However, retrieving from a unified knowledge

researcharxiv-cs-cl
1 Jun 2026
Model Releases

MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents

DGX agent

arXiv:2605.30727v1 Announce Type: new Abstract: Deep research agents increasingly combine private local documents with external tools like web retrieval, creating a privacy risk: an agent's external q

model-releasesarxiv-cs-cl
1 Jun 2026
Agents

Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely

DGX agent

arXiv:2605.31387v1 Announce Type: new Abstract: Robots operating in diverse environments rely on visual input to interpret objects and spatial layouts. In human-collaborative tasks, they are expected

agentsarxiv-cs-cl
1 Jun 2026
Research

Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages

DGX agent

arXiv:2605.31136v1 Announce Type: new Abstract: In automated fact-checking (AFC), check-worthiness detection identifies claims requiring verification based on domain-specific criteria. On Wikipedia, t

researcharxiv-cs-cl
1 Jun 2026
Model Releases

NeUQI: Near-Optimal Uniform Quantization Parameter Initialization for Low-Bit LLMs

DGX agent

arXiv:2505.17595v4 Announce Type: replace-cross Abstract: Large language models (LLMs) achieve impressive performance across domains but face significant challenges when deployed on consumer-grade GPU

model-releasesarxiv-cs-cl
1 Jun 2026
Safety

Neuron-Level Interventions for Gendered and Gender-Neutral Generation in Language Models

DGX agent

arXiv:2605.30717v1 Announce Type: new Abstract: Language models (LMs) can produce gendered language and stereotypes even when given neutral prompts. Most prior work on gender bias in LMs primarily exa

safetyarxiv-cs-cl
1 Jun 2026
Safety

On the 'Induction Bias' in Sequence Models

DGX agent

arXiv:2602.18333v2 Announce Type: replace-cross Abstract: Despite the remarkable practical success of transformer-based language models, recent work has raised concerns about their ability to perform

safetyarxiv-cs-cl
1 Jun 2026
Model Releases

Pairwise Reference Alignment as a Model-Level Ordinal Observable

DGX agent

arXiv:2605.30758v1 Announce Type: new Abstract: Pairwise preference data is widely used in language-model evaluation and alignment, often for model ranking, reward modeling, or preference optimization

model-releasesarxiv-cs-cl
1 Jun 2026
Hardware

ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs

DGX agent

arXiv:2602.07721v3 Announce Type: replace-cross Abstract: KV-cache retrieval is essential for long-context LLM inference, yet existing methods struggle with distribution drift and high latency at scal

hardwarearxiv-cs-cl
1 Jun 2026
Safety

*-PLUIE: Personalisable metric with Llm Used for Improved Evaluation

DGX agent

arXiv:2602.15778v2 Announce Type: replace Abstract: Evaluating the quality of automatically generated text often relies on LLM-as-a-judge (LLM-judge) methods. While effective, these approaches are com

safetyarxiv-cs-cl
1 Jun 2026
Safety

Preference-Aware Rubric Learning for Personalized Evaluation

DGX agent

arXiv:2605.31545v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve from general-purpose assistants to user-centric agents, personalization has become central to aligning model beha

safetyarxiv-cs-cl
1 Jun 2026
Model Releases

Probing the Prompt KV Cache: Where It Becomes Dispensable

DGX agent

arXiv:2605.30574v1 Announce Type: new Abstract: Prior KV cache compression schemes empirically demonstrate that the prompt cache is partially redundant during decoding, dropping or summarising entries

model-releasesarxiv-cs-cl
1 Jun 2026
Research

Protocol for evaluating ChatGPT in biomedical association generation and verification using a RAG-enabled, cross-model majority voting workflow

DGX agent

arXiv:2605.30400v1 Announce Type: new Abstract: We present a protocol to evaluate ChatGPT's ability to generate disease-centric biomedical associations. It outlines how we generate the associations, v

researcharxiv-cs-cl
1 Jun 2026
Model Releases

Query-focused and Memory-aware Reranker for Long Context Processing

DGX agent

arXiv:2602.12192v3 Announce Type: replace Abstract: Built upon the existing analysis of retrieval heads in large language models, we propose an alternative reranking framework that trains models to es

model-releasesarxiv-cs-cl
1 Jun 2026
Safety

Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses

DGX agent

arXiv:2504.11972v3 Announce Type: replace Abstract: Extractive QA tasks are commonly evaluated using Exact Match (EM) and F1-score, but these metrics often fail to reflect true model performance. Rece

safetyarxiv-cs-cl
1 Jun 2026
Research

Refining Word-Based Grammatical Error Annotation for L2 Korean

DGX agent

arXiv:2605.30545v1 Announce Type: new Abstract: Korean grammatical error correction (K-GEC) presents a structural mismatch between word-based evaluation and the morpheme-level locus of many learner er

researcharxiv-cs-cl
1 Jun 2026
Safety

Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards

DGX agent

arXiv:2605.31328v1 Announce Type: new Abstract: Emergent misalignment (EM) is the surprising tendency of language models to become broadly misaligned after fine-tuning on narrowly misaligned examples.

safetyarxiv-cs-cl
1 Jun 2026
Applications

Reliable Multilingual Orthopedic Decision Support from Clinical Narratives: Language-Aware Adaptation and Verification-Guided Deferral

DGX agent

arXiv:2605.31512v1 Announce Type: new Abstract: Multilingual orthopedic decision support remains challenging in low-resource healthcare settings, where clinical narratives contain specialized terminol

applicationsarxiv-cs-cl
1 Jun 2026
Research

Rethinking Sparse Mixture of Experts from a Unified Perspective

DGX agent

arXiv:2503.22996v3 Announce Type: replace Abstract: Sparse Mixture of Experts (SMoE) models scale the capacity of models while maintaining constant computational overhead. SMoE methods fall into two c

researcharxiv-cs-cl
1 Jun 2026
Model Releases

Scaling Multi-Hop Training Data via Graph-Constrained Path Selection

DGX agent

arXiv:2605.31238v1 Announce Type: new Abstract: Endowing large language models with compositional reasoning over specialized documents requires multi-hop training data at scale, where such data rarely

model-releasesarxiv-cs-cl
1 Jun 2026
Research

SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks

DGX agent

arXiv:2605.31433v1 Announce Type: new Abstract: Self-play can train language models without external supervision. However, existing methods require rule-checkable answers, leaving open-ended tasks dep

researcharxiv-cs-cl
1 Jun 2026
Research

Self-Reflective Generation at Test Time

DGX agent

arXiv:2510.02919v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly solve complex reasoning tasks via long chain-of-thought, but their forward-only autoregressive generation

researcharxiv-cs-cl
1 Jun 2026
Safety

Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures

DGX agent

arXiv:2605.30608v1 Announce Type: new Abstract: Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains ch

safetyarxiv-cs-cl
1 Jun 2026
Research

Semantic Triplet Restoration: A Novel Protocol for Hierarchical Table Understanding in Large Language Models

DGX agent

arXiv:2605.31550v1 Announce Type: new Abstract: Table question answering requires models to recover semantic relations encoded implicitly by two-dimensional layout, merged cells, and hierarchical head

researcharxiv-cs-cl
1 Jun 2026
Model Releases

SERA: Soft-Verified Efficient Repository Agents

DGX agent

arXiv:2601.20789v3 Announce Type: replace Abstract: Open-weight coding agents should hold a fundamental advantage over closed-source systems because they can specialize to private codebases, encoding

model-releasesarxiv-cs-cl
1 Jun 2026
Safety

Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents

DGX agent

arXiv:2605.30723v1 Announce Type: new Abstract: LLM agents increasingly retrieve externally curated skills-procedural instructions retrieved at decision time-to improve performance on long-horizon int

safetyarxiv-cs-cl
1 Jun 2026
Research

Speculative Decoding Across Languages

DGX agent

arXiv:2605.30580v1 Announce Type: new Abstract: Speculative decoding has become a crucial component of large language model (LLM) inference, enabling faster generation by drafting multiple tokens and

researcharxiv-cs-cl
1 Jun 2026
Research

Speculative Pipeline Decoding: Higher-Accruacy and Zero-Bubble Speculation via Pipeline Parallelism

DGX agent

arXiv:2605.30852v1 Announce Type: new Abstract: Speculative Decoding (SD) accelerates low-concurrency LLM inference by employing a draft-then-verify paradigm. However, mainstream methods typically rel

researcharxiv-cs-cl
1 Jun 2026
Safety

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation

DGX agent

arXiv:2511.11440v3 Announce Type: replace-cross Abstract: Performance gains of Vision Language Models (VLMs) obtained by fine-tuning are generally based on ad hoc data collection and annotation of rea

safetyarxiv-cs-cl
1 Jun 2026
Model Releases

TaxoBell: Gaussian Box Embeddings for Self-Supervised Taxonomy Expansion

DGX agent

arXiv:2601.09633v2 Announce Type: replace Abstract: Taxonomies form the backbone of structured knowledge representation across diverse domains, enabling applications such as e-commerce and semantic se

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

DGX agent

arXiv:2605.30673v1 Announce Type: new Abstract: Classroom videos contain observable teaching practices, but their pedagogical and visual signals are rarely organized in forms suitable for model evalua

model-releasesarxiv-cs-cl
1 Jun 2026
Safety

The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement

DGX agent

arXiv:2605.30888v1 Announce Type: new Abstract: Building strong reward models (RMs) for language model alignment is bottlenecked by the cost and difficulty of acquiring diverse and reliable preference

safetyarxiv-cs-cl
1 Jun 2026
Research

The Latin Substrate: How Language Models Represent and Mediate Script Choice

DGX agent

arXiv:2605.31363v1 Announce Type: new Abstract: Many languages are written in multiple scripts, requiring large language models (LLMs) to generate equivalent linguistic content in distinct orthographi

researcharxiv-cs-cl
1 Jun 2026
Tutorials

The relative strength of hierarchical structure and statistics differs across the measures in naturalistic reading

DGX agent

arXiv:2509.23195v2 Announce Type: replace Abstract: The hierarchical syntactic structure and non-hierarchical, statistical, or sequential factors have long been framed as rival theories in accounting

tutorialsarxiv-cs-cl
1 Jun 2026
Applications

Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining

DGX agent

arXiv:2605.31069v1 Announce Type: cross Abstract: Accurately predicting future events is fundamental to content understanding and decision-making across various domains. While prior research has prima

applicationsarxiv-cs-cl
1 Jun 2026
Research

Towards Efficient LLMs Annealing with Principled Sample Selection

DGX agent

arXiv:2605.31175v1 Announce Type: new Abstract: The annealing phase is a pivotal convergence stage in LLM pre-training that ultimately determines final model quality. However, effectively selecting tr

researcharxiv-cs-cl
1 Jun 2026
Model Releases

TRACE: Discovering Task-Specific Parameter via Adaptation-Aware Probing for Continual Fine-Tuning

DGX agent

arXiv:2605.31025v1 Announce Type: new Abstract: In real-world deployment, LLMs are often adapted continually across tasks to keep LLMs up-to-date in production, where new fine-tuning should preserve p

model-releasesarxiv-cs-cl
1 Jun 2026
Safety

Traceable by Design: An LLM Pipeline and Dashboard for EU Regulatory Consultation Analysis

DGX agent

arXiv:2605.30995v1 Announce Type: cross Abstract: Public consultations generate large volumes of data in the form of stakeholder submissions that are practically unfeasible to analyse manually. We pre

safetyarxiv-cs-cl
1 Jun 2026
Tutorials

Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing

DGX agent

arXiv:2605.31367v1 Announce Type: cross Abstract: Token mixing layers play a key role in how language models can learn and generate long-range dependencies. Their efficiency relies on the necessary tr

tutorialsarxiv-cs-cl
1 Jun 2026
Model Releases

Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation Workflows

DGX agent

arXiv:2605.31452v1 Announce Type: new Abstract: Building on our previous work, this paper develops practical, low-barrier methods for freelance translators and smaller language service providers to ev

model-releasesarxiv-cs-cl
1 Jun 2026
Research

TransLPRNet: Lite Vision-Language Network for Single/Dual-line Chinese License Plate Recognition

DGX agent

arXiv:2507.17335v2 Announce Type: replace-cross Abstract: License plate recognition in open environments is widely applicable across various domains; however, the diversity of license plate types and

researcharxiv-cs-cl
1 Jun 2026
Model Releases

Triaging Threats to Specialized Guardrails

DGX agent

arXiv:2605.30693v1 Announce Type: cross Abstract: Building robust safety guardrails is essential for deploying Large Language Models across diverse real-world applications. However, this goal remains

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

TSM-Bench: Detecting LLM-Generated Text in Real-World Wikipedia Editing Practices

DGX agent

arXiv:2605.31113v1 Announce Type: new Abstract: Automatically detecting machine-generated text (MGT) is critical to maintaining the knowledge integrity of user-generated content (UGC) platforms such a

model-releasesarxiv-cs-cl
1 Jun 2026
← Previous
1…6566676869…161
Next →