AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Research

'Intelegi Romaneste?'' A Recipe for Romanian Vision-Language Models

DGX agent

arXiv:2605.31401v1 Announce Type: new Abstract: Vision-Language Models (VLMs) largely follow the text-only LLM trajectory, excelling on English benchmarks but sharply degrading on low-resource languag

researcharxiv-cs-cl
1 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Knowledge Boundary Probing and Demand-Guided Intervention for LLM-Based Power System Code Generation

DGX agent

arXiv:2605.31478v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to automate power-system analysis, but many utilities and energy-research labs require on-premise s

model-releasesarxiv-cs-cl
1 Jun 2026
Research

Knowledge Graph-Enhanced Zero-Shot Topic Classification: A Multi-Strategy Comparative Study

DGX agent

arXiv:2605.30465v1 Announce Type: new Abstract: Multi-label topic classification without labeled training data is a challenging task, specially when documents contain complex relational information. W

researcharxiv-cs-cl
1 Jun 2026
Research

Language Models Can Resolve Reference Compositionally, But It's Not Their Native Strength: The Case of the Personal Relation Task

DGX agent

arXiv:2605.31480v1 Announce Type: new Abstract: Do neural models, such as Large Language Models, genuinely acquire compositional abilities for interpretation of natural language? When we talk about se

researcharxiv-cs-cl
1 Jun 2026
Research

Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization

DGX agent

arXiv:2605.31312v1 Announce Type: cross Abstract: Multimodal hallucination remains a persistent challenge for Vision-Language Models (VLMs). Standard textual Direct Preference Optimization (DPO) often

researcharxiv-cs-cl
1 Jun 2026
Model Releases

Learning Whom to Trust: Market-Feedback Adaptive Retrieval for Frozen LLMs in Event-Driven Financial RAG

DGX agent

arXiv:2605.31201v1 Announce Type: new Abstract: Financial retrieval-augmented generation (RAG) systems typically rank evidence by textual relevance, but in financial markets the useful evidence source

model-releasesarxiv-cs-cl
1 Jun 2026
Research

Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs

DGX agent

arXiv:2605.30501v1 Announce Type: new Abstract: Watermarking embeds statistical signatures in AI-generated text for detection and attribution. We reveal a fundamental vulnerability: when users access

researcharxiv-cs-cl
1 Jun 2026
Agents

LLM Anonymization Against Agentic Re-Identificatio

DGX agent

arXiv:2605.30848v1 Announce Type: cross Abstract: Agentic LLMs with web search change the threat model for text anonymization: weak contextual cues can become cross-referenceable evidence for re-ident

agentsarxiv-cs-cl
1 Jun 2026
Safety

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories

DGX agent

arXiv:2605.31381v1 Announce Type: new Abstract: We evaluate the consistency of automated judges in conducting a multi-dimensional safety evaluation in a reference-free setup. Our results indicate that

safetyarxiv-cs-cl
1 Jun 2026
Local Ai

LocalSUG: City-Preference-Enhanced LLM for Query Suggestion in Local-Life Services

DGX agent

arXiv:2603.04946v2 Announce Type: replace Abstract: In local-life service platforms, query suggestion reduces user effort by generating candidate queries from input prefixes. Traditional multi-stage s

local-aiarxiv-cs-cl
1 Jun 2026
Model Releases

MAAT: Multi-phase Adapter-Aware Targeted Unlearning

DGX agent

arXiv:2605.30514v1 Announce Type: cross Abstract: Machine unlearning evaluation is structurally skewed: Why-type questions, which probe causal and relational knowledge, comprise less than 0.06% of Cou

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

MADS: Model-Aware Diverse Core Set Selection for Instruction Tuning

DGX agent

arXiv:2605.30857v1 Announce Type: new Abstract: Instruction fine-tuning is employed to enhance the instruction-following ability of large language models (LLMs). As the amount of instruction fine-tuni

model-releasesarxiv-cs-cl
1 Jun 2026
Local Ai

Measuring, Localizing, and Ablating Alignment Signatures in LLMs

DGX agent

arXiv:2605.30526v1 Announce Type: cross Abstract: Aligned language models often exhibit a recognizable AI-like style, yet its connection to post-training and internal representations remains poorly un

local-aiarxiv-cs-cl
1 Jun 2026
Model Releases

Mellum2 Technical Report

DGX agent

arXiv:2605.31268v1 Announce Type: new Abstract: We present Mellum 2, an open-weight 12B-parameter Mixture-of-Experts (MoE) language model with 2.5B active parameters per token. Mellum 2 is a general-p

model-releasesarxiv-cs-cl
1 Jun 2026
Local Ai

Memory-Efficient Structured Backpropagation for On-Device LLM Fine-Tuning

DGX agent

arXiv:2602.13069v2 Announce Type: replace-cross Abstract: On-device fine-tuning enables privacy-preserving personalization of large language models, but mobile devices impose severe memory constraints

local-aiarxiv-cs-cl
1 Jun 2026
Model Releases

MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft

DGX agent

arXiv:2605.30931v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to susta

model-releasesarxiv-cs-cl
1 Jun 2026
Research

MoG: Mixture of Experts for Graph-based Retrieval-Augmented Generation

DGX agent

arXiv:2605.31010v1 Announce Type: new Abstract: Retrieval-augmented generation is intensively studied to ground large language models on external evidence. However, retrieving from a unified knowledge

researcharxiv-cs-cl
1 Jun 2026
Model Releases

MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents

DGX agent

arXiv:2605.30727v1 Announce Type: new Abstract: Deep research agents increasingly combine private local documents with external tools like web retrieval, creating a privacy risk: an agent's external q

model-releasesarxiv-cs-cl
1 Jun 2026
Agents

Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely

DGX agent

arXiv:2605.31387v1 Announce Type: new Abstract: Robots operating in diverse environments rely on visual input to interpret objects and spatial layouts. In human-collaborative tasks, they are expected

agentsarxiv-cs-cl
1 Jun 2026
Research

Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages

DGX agent

arXiv:2605.31136v1 Announce Type: new Abstract: In automated fact-checking (AFC), check-worthiness detection identifies claims requiring verification based on domain-specific criteria. On Wikipedia, t

researcharxiv-cs-cl
1 Jun 2026
Model Releases

NeUQI: Near-Optimal Uniform Quantization Parameter Initialization for Low-Bit LLMs

DGX agent

arXiv:2505.17595v4 Announce Type: replace-cross Abstract: Large language models (LLMs) achieve impressive performance across domains but face significant challenges when deployed on consumer-grade GPU

model-releasesarxiv-cs-cl
1 Jun 2026
Safety

Neuron-Level Interventions for Gendered and Gender-Neutral Generation in Language Models

DGX agent

arXiv:2605.30717v1 Announce Type: new Abstract: Language models (LMs) can produce gendered language and stereotypes even when given neutral prompts. Most prior work on gender bias in LMs primarily exa

safetyarxiv-cs-cl
1 Jun 2026
Safety

On the 'Induction Bias' in Sequence Models

DGX agent

arXiv:2602.18333v2 Announce Type: replace-cross Abstract: Despite the remarkable practical success of transformer-based language models, recent work has raised concerns about their ability to perform

safetyarxiv-cs-cl
1 Jun 2026
Model Releases

Pairwise Reference Alignment as a Model-Level Ordinal Observable

DGX agent

arXiv:2605.30758v1 Announce Type: new Abstract: Pairwise preference data is widely used in language-model evaluation and alignment, often for model ranking, reward modeling, or preference optimization

model-releasesarxiv-cs-cl
1 Jun 2026
Hardware

ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs

DGX agent

arXiv:2602.07721v3 Announce Type: replace-cross Abstract: KV-cache retrieval is essential for long-context LLM inference, yet existing methods struggle with distribution drift and high latency at scal

hardwarearxiv-cs-cl
1 Jun 2026
Safety

*-PLUIE: Personalisable metric with Llm Used for Improved Evaluation

DGX agent

arXiv:2602.15778v2 Announce Type: replace Abstract: Evaluating the quality of automatically generated text often relies on LLM-as-a-judge (LLM-judge) methods. While effective, these approaches are com

safetyarxiv-cs-cl
1 Jun 2026
Safety

Preference-Aware Rubric Learning for Personalized Evaluation

DGX agent

arXiv:2605.31545v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve from general-purpose assistants to user-centric agents, personalization has become central to aligning model beha

safetyarxiv-cs-cl
1 Jun 2026
Model Releases

Probing the Prompt KV Cache: Where It Becomes Dispensable

DGX agent

arXiv:2605.30574v1 Announce Type: new Abstract: Prior KV cache compression schemes empirically demonstrate that the prompt cache is partially redundant during decoding, dropping or summarising entries

model-releasesarxiv-cs-cl
1 Jun 2026
Research

Protocol for evaluating ChatGPT in biomedical association generation and verification using a RAG-enabled, cross-model majority voting workflow

DGX agent

arXiv:2605.30400v1 Announce Type: new Abstract: We present a protocol to evaluate ChatGPT's ability to generate disease-centric biomedical associations. It outlines how we generate the associations, v

researcharxiv-cs-cl
1 Jun 2026
Model Releases

Query-focused and Memory-aware Reranker for Long Context Processing

DGX agent

arXiv:2602.12192v3 Announce Type: replace Abstract: Built upon the existing analysis of retrieval heads in large language models, we propose an alternative reranking framework that trains models to es

model-releasesarxiv-cs-cl
1 Jun 2026
Safety

Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses

DGX agent

arXiv:2504.11972v3 Announce Type: replace Abstract: Extractive QA tasks are commonly evaluated using Exact Match (EM) and F1-score, but these metrics often fail to reflect true model performance. Rece

safetyarxiv-cs-cl
1 Jun 2026
Research

Refining Word-Based Grammatical Error Annotation for L2 Korean

DGX agent

arXiv:2605.30545v1 Announce Type: new Abstract: Korean grammatical error correction (K-GEC) presents a structural mismatch between word-based evaluation and the morpheme-level locus of many learner er

researcharxiv-cs-cl
1 Jun 2026
Safety

Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards

DGX agent

arXiv:2605.31328v1 Announce Type: new Abstract: Emergent misalignment (EM) is the surprising tendency of language models to become broadly misaligned after fine-tuning on narrowly misaligned examples.

safetyarxiv-cs-cl
1 Jun 2026
Applications

Reliable Multilingual Orthopedic Decision Support from Clinical Narratives: Language-Aware Adaptation and Verification-Guided Deferral

DGX agent

arXiv:2605.31512v1 Announce Type: new Abstract: Multilingual orthopedic decision support remains challenging in low-resource healthcare settings, where clinical narratives contain specialized terminol

applicationsarxiv-cs-cl
1 Jun 2026
Research

Rethinking Sparse Mixture of Experts from a Unified Perspective

DGX agent

arXiv:2503.22996v3 Announce Type: replace Abstract: Sparse Mixture of Experts (SMoE) models scale the capacity of models while maintaining constant computational overhead. SMoE methods fall into two c

researcharxiv-cs-cl
1 Jun 2026
Model Releases

Scaling Multi-Hop Training Data via Graph-Constrained Path Selection

DGX agent

arXiv:2605.31238v1 Announce Type: new Abstract: Endowing large language models with compositional reasoning over specialized documents requires multi-hop training data at scale, where such data rarely

model-releasesarxiv-cs-cl
1 Jun 2026
Research

SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks

DGX agent

arXiv:2605.31433v1 Announce Type: new Abstract: Self-play can train language models without external supervision. However, existing methods require rule-checkable answers, leaving open-ended tasks dep

researcharxiv-cs-cl
1 Jun 2026
Research

Self-Reflective Generation at Test Time

DGX agent

arXiv:2510.02919v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly solve complex reasoning tasks via long chain-of-thought, but their forward-only autoregressive generation

researcharxiv-cs-cl
1 Jun 2026
Safety

Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures

DGX agent

arXiv:2605.30608v1 Announce Type: new Abstract: Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains ch

safetyarxiv-cs-cl
1 Jun 2026
Research

Semantic Triplet Restoration: A Novel Protocol for Hierarchical Table Understanding in Large Language Models

DGX agent

arXiv:2605.31550v1 Announce Type: new Abstract: Table question answering requires models to recover semantic relations encoded implicitly by two-dimensional layout, merged cells, and hierarchical head

researcharxiv-cs-cl
1 Jun 2026
Model Releases

SERA: Soft-Verified Efficient Repository Agents

DGX agent

arXiv:2601.20789v3 Announce Type: replace Abstract: Open-weight coding agents should hold a fundamental advantage over closed-source systems because they can specialize to private codebases, encoding

model-releasesarxiv-cs-cl
1 Jun 2026
Safety

Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents

DGX agent

arXiv:2605.30723v1 Announce Type: new Abstract: LLM agents increasingly retrieve externally curated skills-procedural instructions retrieved at decision time-to improve performance on long-horizon int

safetyarxiv-cs-cl
1 Jun 2026
Research

Speculative Decoding Across Languages

DGX agent

arXiv:2605.30580v1 Announce Type: new Abstract: Speculative decoding has become a crucial component of large language model (LLM) inference, enabling faster generation by drafting multiple tokens and

researcharxiv-cs-cl
1 Jun 2026
Research

Speculative Pipeline Decoding: Higher-Accruacy and Zero-Bubble Speculation via Pipeline Parallelism

DGX agent

arXiv:2605.30852v1 Announce Type: new Abstract: Speculative Decoding (SD) accelerates low-concurrency LLM inference by employing a draft-then-verify paradigm. However, mainstream methods typically rel

researcharxiv-cs-cl
1 Jun 2026
Safety

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation

DGX agent

arXiv:2511.11440v3 Announce Type: replace-cross Abstract: Performance gains of Vision Language Models (VLMs) obtained by fine-tuning are generally based on ad hoc data collection and annotation of rea

safetyarxiv-cs-cl
1 Jun 2026
Model Releases

TaxoBell: Gaussian Box Embeddings for Self-Supervised Taxonomy Expansion

DGX agent

arXiv:2601.09633v2 Announce Type: replace Abstract: Taxonomies form the backbone of structured knowledge representation across diverse domains, enabling applications such as e-commerce and semantic se

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

DGX agent

arXiv:2605.30673v1 Announce Type: new Abstract: Classroom videos contain observable teaching practices, but their pedagogical and visual signals are rarely organized in forms suitable for model evalua

model-releasesarxiv-cs-cl
1 Jun 2026
Safety

The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement

DGX agent

arXiv:2605.30888v1 Announce Type: new Abstract: Building strong reward models (RMs) for language model alignment is bottlenecked by the cost and difficulty of acquiring diverse and reliable preference

safetyarxiv-cs-cl
1 Jun 2026
← Previous
1…6667686970…162
Next →