AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
12 May 2026

Decomposing and Steering Functional Metacognition in Large Language Models

Model ReleasesDGX agent

arXiv:2605.08942v1 Announce Type: new Abstract: Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, often adapting their reasoning strategies

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification

ResearchDGX agent

arXiv:2605.09269v1 Announce Type: new Abstract: Aligning Multimodal Large Language Models (MLLMs) requires reliable reward models, yet existing single-step evaluators can suffer from lazy judging, exp

Deterministic Differentiable Structured Pruning for Large Language Models

ResearchDGX agent

arXiv:2603.08065v2 Announce Type: replace-cross Abstract: Structured pruning reduces LLM inference cost by removing low-importance architectural components. This can be viewed as learning a multiplica


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

DGPO: Beyond Pairwise Preferences with Directional Consistent Groupwise Optimization

SafetyDGX agent

arXiv:2605.10863v1 Announce Type: new Abstract: Although Large Language Models (LLMs) have made remarkable progress, current preference optimization methods still struggle to align directional consist

Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling

TutorialsDGX agent

arXiv:2605.08477v1 Announce Type: new Abstract: Explicit planning is a critical capability for LLM-based agents solving complex data-centric tasks, which require precise tool calling over external dat

DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding

Model ReleasesDGX agent

arXiv:2605.08888v1 Announce Type: new Abstract: Evaluating whether Multimodal Large Language Models can produce trustworthy, verifiable reasoning over long, visually rich documents requires evaluation

Dolphin-CN-Dialect: Where Chinese Dialects Matter

ApplicationsDGX agent

arXiv:2605.08961v1 Announce Type: new Abstract: We present Dolphin-CN-Dialect, a streaming-capable ASR model with a focus on Chinese and dialect-rich scenarios. Compared to the previous version, Dolph

Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval

Model ReleasesDGX agent

arXiv:2504.21015v4 Announce Type: replace-cross Abstract: Training effective dense retrieval models typically relies on hard negative (HN) examples mined from large document corpora using methods such

Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training

TutorialsDGX agent

arXiv:2603.04415v2 Announce Type: replace Abstract: Reasoning post-training improves Large Language Models (LLMs) on complex tasks such as mathematics and coding, but its benefits across diverse multi

Dynamic Meta-Metrics: Source-Sentence Conditioned Weighting for MT Evaluation

ResearchDGX agent

arXiv:2605.09098v1 Announce Type: new Abstract: We propose Dynamic Meta-Metrics (DMM), a framework for machine translation evaluation that learns source-sentence conditioned combinations of existing m

Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2605.10923v1 Announce Type: cross Abstract: Large language model agents increasingly rely on external skills to solve complex tasks, where skills act as modular units that extend their capabilit

EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments

Model ReleasesDGX agent

arXiv:2506.08136v3 Announce Type: replace Abstract: We introduce EconWebArena, a benchmark for evaluating autonomous agents on complex, multimodal economic tasks in realistic web environments. The ben

EdgeFlowerTune: Evaluating Federated LLM Fine-Tuning Under Realistic Edge System Constraints

Model ReleasesDGX agent

arXiv:2605.08636v1 Announce Type: new Abstract: Federated fine-tuning offers a promising paradigm for adapting large language models (LLMs) on edge devices by leveraging the rich, diverse, and continu

Edit-Based Refinement for Parallel Masked Diffusion Language Models

ResearchDGX agent

arXiv:2605.09603v1 Announce Type: new Abstract: Masked diffusion language models enable parallel token generation and offer improved decoding efficiency over autoregressive models. However, their perf

EMO: Pretraining Mixture of Experts for Emergent Modularity

ResearchDGX agent

arXiv:2605.06663v2 Announce Type: replace Abstract: Large language models are typically deployed as monolithic systems, requiring the full model even when applications need only a narrow subset of cap

EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding

Model ReleasesDGX agent

arXiv:2605.08847v1 Announce Type: new Abstract: In the context of today's high-pressure, aging society, the demand for large-scale emotional models capable of providing empathetic support is more crit

ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room

Model ReleasesDGX agent

arXiv:2505.22919v3 Announce Type: replace Abstract: Existing benchmarks for evaluating the clinical reasoning capabilities of large language models (LLMs) often lack a clear definition of 'clinical re

Evaluating Pragmatic Reasoning in Large Language Models: Evidence from Scalar Diversity

ResearchDGX agent

arXiv:2605.09042v1 Announce Type: new Abstract: Evaluating pragmatic reasoning in large language models (LLMs) remains challenging because model behavior can vary depending on evaluation methods. Prev

Evolving Knowledge Distillation for Lightweight Neural Machine Translation

ResearchDGX agent

arXiv:2605.09924v1 Announce Type: new Abstract: Recent advancements in Neural Machine Translation (NMT) have significantly improved translation quality. However, the increasing size and complexity of

Extending Confidence-Based Text2Cypher with Grammar and Schema Aware Filtering

ResearchDGX agent

arXiv:2605.10318v1 Announce Type: new Abstract: Large language models (LLMs) allow users to query databases using natural language by translating questions into executable queries. Despite strong prog

Fast-MIA: Efficient and Scalable Membership Inference for LLMs

ResearchDGX agent

arXiv:2510.23074v2 Announce Type: replace-cross Abstract: We propose Fast-MIA (https://github.com/Nikkei/fast-mia), a Python library for efficiently evaluating membership inference attacks (MIA) again

Feature Rivalry in Sparse Autoencoder Representations: A Mechanistic Study of Uncertainty-Driven Feature Competition in LLMs

Model ReleasesDGX agent

arXiv:2605.08149v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) decompose large language model representations into interpretable features, but how these features interact under uncertain

Federated Language Models Under Bandwidth Budgets: Distillation Rates and Conformal Coverage

Model ReleasesDGX agent

arXiv:2605.09986v1 Announce Type: cross Abstract: Training a language model on data scattered across bandwidth-limited nodes that cannot be centralized is a setting that arises in clinical networks, e

FERA: Uncertainty-Aware Federated Reasoning for Large Language Models

ResearchDGX agent

arXiv:2605.10082v1 Announce Type: new Abstract: Large language models (LLMs) exhibit strong reasoning capabilities when guided by high-quality demonstrations, yet such data is often distributed across

Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain

Model ReleasesDGX agent

arXiv:2605.09106v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in financial contexts, raising critical concerns about reliability, alignment, and susceptibility

FinMoji: A Framework for Emoji-driven Sentiment Analysis in Financial Social Media

ResearchDGX agent

arXiv:2605.09469v1 Announce Type: new Abstract: This paper explores the use of emojis in financial sentiment analysis, focusing on the social media platform StockTwits. Emojis, increasingly prevalent

First, Do No Harm: AI Supervisor Scaffolds Novice Growth in Counselor Education

TutorialsDGX agent

arXiv:2508.09042v3 Announce Type: replace Abstract: The most dangerous mistakes a novice counselor makes are not the obvious ones: they are utterances that sound caring while quietly violating profess

FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning

AgentsDGX agent

arXiv:2605.09932v1 Announce Type: new Abstract: Large language models can now process increasingly long inputs, yet their ability to effectively use information spread across long contexts remains lim

Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling

Model ReleasesDGX agent

arXiv:2512.02010v5 Announce Type: replace Abstract: As large language models have grown larger, interest has grown in low-precision numerical formats such as NVFP4 as a way to improve speed and reduce

Frame In, Frame Out: Measuring Framing Bias in LLM-Generated News Summaries

Model ReleasesDGX agent

arXiv:2505.05406v2 Announce Type: replace Abstract: News headlines and summaries shape how events are interpreted through selective emphasis and omission, a phenomenon commonly referred to as framing.

GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM Collectives

Model ReleasesDGX agent

arXiv:2605.09027v1 Announce Type: new Abstract: In multi-agent systems (MAS), a single deceptive agent can nullify all gains of an agentic AI collective and evade deployed defenses. However, existing

GIFT: Guided Importance-Aware Fine-Tuning for Diffusion Language Models

Model ReleasesDGX agent

arXiv:2509.20863v3 Announce Type: replace Abstract: Diffusion models have recently shown strong potential in language modeling, offering faster generation compared to traditional autoregressive approa

GLiNER-Relex: A Unified Framework for Joint Named Entity Recognition and Relation Extraction

Model ReleasesDGX agent

arXiv:2605.10108v1 Announce Type: new Abstract: Joint named entity recognition (NER) and relation extraction (RE) is a fundamental task in natural language processing for constructing knowledge graphs

GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression

AgentsDGX agent

arXiv:2605.09100v1 Announce Type: new Abstract: Text embedding and generative tasks are usually trained separately based on large language models (LLMs) nowadays. This causes a large amount of trainin

Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking

Model ReleasesDGX agent

arXiv:2605.10893v1 Announce Type: new Abstract: Large vision-language models suffer from visual ungroundedness: they can produce a fluent, confident, and even correct response driven entirely by langu

Grounded Satirical Generation with RAG

ResearchDGX agent

arXiv:2605.10853v1 Announce Type: new Abstract: Humor generation remains challenging task for Large Language Models (LLMs), due to their subjective nature. We focus on satire, a form of humor strongly

Hint Tuning: Less Data Makes Better Reasoners

Model ReleasesDGX agent

arXiv:2605.08665v1 Announce Type: new Abstract: Large reasoning models achieve high accuracy through extended chain-of-thought but generate 5--8 more tokens than necessary, applying verbose reasoning

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models

Model ReleasesDGX agent

arXiv:2404.18923v5 Announce Type: replace Abstract: We introduce Holmes, a new benchmark designed to assess language models (LMs) linguistic competence - their unconscious understanding of linguistic

How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits

ResearchDGX agent

arXiv:2605.08348v1 Announce Type: new Abstract: The circuits framework in mechanistic interpretability aims to identify causally important sparse subgraphs of model components, typically evaluated by

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue

ResearchDGX agent

arXiv:2605.10199v1 Announce Type: new Abstract: Full-duplex spoken dialogue requires a model to keep listening while generating its own spoken response. This is challenging for large language models (

ICT-NLP at SemEval-2026 Task 3: Less Is More -- Multilingual Encoder with Joint Training and Adaptive Ensemble for Dimensional Aspect Sentiment Regression

ResearchDGX agent

arXiv:2605.10560v1 Announce Type: new Abstract: This paper describes our system to SemEval-2026 Task 3 Track A Subtask 1 on Dimensional Aspect Sentiment Regression (DimASR). We propose a lightweight a

Incremental Multilingual Text2Cypher with Adapter Combination

Model ReleasesDGX agent

arXiv:2601.16097v2 Announce Type: replace Abstract: Large Language Models enable users to access database using natural language interfaces using tools like Text2SQL, Text2SPARQL, and Text2Cypher, whi

Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables

Model ReleasesDGX agent

arXiv:2605.10039v1 Announce Type: cross Abstract: Frontier coding agents read configuration files (CLAUDE.md, AGENTS.md, Cursor Rules) at session start and are expected to follow the conventions insid

Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration

SafetyDGX agent

arXiv:2602.03677v2 Announce Type: replace Abstract: Modality following is the ability to selectively leverage multimodal contexts based on user instructions. It is fundamental to the safety and reliab

jina-embeddings-v5-omni: Text-Geometry-Preserving Multimodal Embeddings via Frozen-Tower Composition

Model ReleasesDGX agent

arXiv:2605.08384v1 Announce Type: new Abstract: In this work, we introduce frozen-encoder model composition, a novel approach to multimodal embedding models. We build on the VLM-style architecture, in

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

Model ReleasesDGX agent

arXiv:2605.09635v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in K-12 education, yet existing benchmarks such as C-Eval, CMMLU, GaokaoBench, and EduEval mainly eva

Language-Conditioned Visual Grounding with CLIP Multilingual

ResearchDGX agent

arXiv:2605.09060v1 Announce Type: new Abstract: Multilingual vision-language models exhibit systematic performance gaps across languages, but the mechanism remains ambiguous: cross-language divergence

Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes

Model ReleasesDGX agent

arXiv:2605.09751v1 Announce Type: new Abstract: Trainable input embedding tables are a standard component of modern language models. We ask whether they are actually necessary at the input interface.

Large Language Models as Students Who Think Aloud: Overly Coherent, Verbose, and Confident

Model ReleasesDGX agent

arXiv:2602.01015v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly embedded in AI-based tutoring systems. Can they faithfully model novice reasoning and metacognitive ju

LEAF-SQL: Level-wise Exploration with Adaptive Fine-graining for Text-to-SQL Skeleton Prediction

Model ReleasesDGX agent

arXiv:2605.09295v1 Announce Type: new Abstract: Text-to-SQL translates natural language questions into executable SQL queries, enabling intuitive database access for non-experts. While large language

Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining

Model ReleasesDGX agent

arXiv:2605.10504v1 Announce Type: new Abstract: A causal-decoder block is hierarchical: lower layers build the residual basis that upper layers attend over. We identify a failure mode in GPT pretraini

Learning More from Less: Exploiting Counterfactuals for Data-Efficient Chart Understanding

ResearchDGX agent

arXiv:2605.10855v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated remarkable progress in chart understanding, largely driven by supervised fine-tuning (SFT) on increasing

Learning to Stay Safe: Adaptive Regularization Against Safety Degradation during Fine-Tuning

SafetyDGX agent

arXiv:2602.17546v2 Announce Type: replace Abstract: Instruction-following language models are trained to be helpful and safe, yet their safety behavior can deteriorate under benign fine-tuning and wor

Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants

ResearchDGX agent

arXiv:2508.16070v3 Announce Type: replace Abstract: Approximately 283 million people worldwide live with visual impairments, motivating increasing research into leveraging Visual Language Models (VLMs

Let the Target Select for Itself: Data Selection via Target-Aligned Paths

SafetyDGX agent

arXiv:2605.09404v1 Announce Type: cross Abstract: Targeted data selection aims to identify training samples from a large candidate pool that improve performance on a specific downstream task. Many rec

LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments

Model ReleasesDGX agent

arXiv:2605.10779v1 Announce Type: cross Abstract: The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content s

LLiMba: Sardinian on a Single GPU -- Adapting a 3B Language Model to a Vanishing Romance Language

Model ReleasesDGX agent

arXiv:2605.09015v1 Announce Type: new Abstract: Sardinian, a Romance language with roughly one million speakers, has minimal presence in modern NLP. Commercial services do not support it, and current

LLM Agents Already Know When to Call Tools -- Even Without Reasoning

Model ReleasesDGX agent

arXiv:2605.09252v1 Announce Type: new Abstract: Tool-augmented LLM agents tend to call tools indiscriminately, even when the model can answer directly. Each unnecessary call wastes API fees and latenc

LLMs with in-context learning for Algorithmic Theoretical Physics

Model ReleasesDGX agent

arXiv:2605.08212v1 Announce Type: cross Abstract: There is an increasing number of algorithmic computations in theoretical physics. These, while conceptually simple, can nevertheless be time-consuming

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories

Model ReleasesDGX agent

arXiv:2509.20909v2 Announce Type: replace Abstract: Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, where benchmark contamination can make memor

← Previous
1…7879808182…129
Next →