AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

EdgeFlowerTune: Evaluating Federated LLM Fine-Tuning Under Realistic Edge System Constraints

DGX agent

arXiv:2605.08636v1 Announce Type: new Abstract: Federated fine-tuning offers a promising paradigm for adapting large language models (LLMs) on edge devices by leveraging the rich, diverse, and continu

model-releasesarxiv-cs-cl
12 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Edit-Based Refinement for Parallel Masked Diffusion Language Models

DGX agent

arXiv:2605.09603v1 Announce Type: new Abstract: Masked diffusion language models enable parallel token generation and offer improved decoding efficiency over autoregressive models. However, their perf

researcharxiv-cs-cl
12 May 2026
Research

EMO: Pretraining Mixture of Experts for Emergent Modularity

DGX agent

arXiv:2605.06663v2 Announce Type: replace Abstract: Large language models are typically deployed as monolithic systems, requiring the full model even when applications need only a narrow subset of cap

researcharxiv-cs-cl
12 May 2026
Model Releases

EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding

DGX agent

arXiv:2605.08847v1 Announce Type: new Abstract: In the context of today's high-pressure, aging society, the demand for large-scale emotional models capable of providing empathetic support is more crit

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room

DGX agent

arXiv:2505.22919v3 Announce Type: replace Abstract: Existing benchmarks for evaluating the clinical reasoning capabilities of large language models (LLMs) often lack a clear definition of 'clinical re

model-releasesarxiv-cs-cl
12 May 2026
Research

Evaluating Pragmatic Reasoning in Large Language Models: Evidence from Scalar Diversity

DGX agent

arXiv:2605.09042v1 Announce Type: new Abstract: Evaluating pragmatic reasoning in large language models (LLMs) remains challenging because model behavior can vary depending on evaluation methods. Prev

researcharxiv-cs-cl
12 May 2026
Research

Evolving Knowledge Distillation for Lightweight Neural Machine Translation

DGX agent

arXiv:2605.09924v1 Announce Type: new Abstract: Recent advancements in Neural Machine Translation (NMT) have significantly improved translation quality. However, the increasing size and complexity of

researcharxiv-cs-cl
12 May 2026
Research

Extending Confidence-Based Text2Cypher with Grammar and Schema Aware Filtering

DGX agent

arXiv:2605.10318v1 Announce Type: new Abstract: Large language models (LLMs) allow users to query databases using natural language by translating questions into executable queries. Despite strong prog

researcharxiv-cs-cl
12 May 2026
Research

Fast-MIA: Efficient and Scalable Membership Inference for LLMs

DGX agent

arXiv:2510.23074v2 Announce Type: replace-cross Abstract: We propose Fast-MIA (https://github.com/Nikkei/fast-mia), a Python library for efficiently evaluating membership inference attacks (MIA) again

researcharxiv-cs-cl
12 May 2026
Model Releases

Feature Rivalry in Sparse Autoencoder Representations: A Mechanistic Study of Uncertainty-Driven Feature Competition in LLMs

DGX agent

arXiv:2605.08149v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) decompose large language model representations into interpretable features, but how these features interact under uncertain

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Federated Language Models Under Bandwidth Budgets: Distillation Rates and Conformal Coverage

DGX agent

arXiv:2605.09986v1 Announce Type: cross Abstract: Training a language model on data scattered across bandwidth-limited nodes that cannot be centralized is a setting that arises in clinical networks, e

model-releasesarxiv-cs-cl
12 May 2026
Research

FERA: Uncertainty-Aware Federated Reasoning for Large Language Models

DGX agent

arXiv:2605.10082v1 Announce Type: new Abstract: Large language models (LLMs) exhibit strong reasoning capabilities when guided by high-quality demonstrations, yet such data is often distributed across

researcharxiv-cs-cl
12 May 2026
Model Releases

Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain

DGX agent

arXiv:2605.09106v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in financial contexts, raising critical concerns about reliability, alignment, and susceptibility

model-releasesarxiv-cs-cl
12 May 2026
Research

FinMoji: A Framework for Emoji-driven Sentiment Analysis in Financial Social Media

DGX agent

arXiv:2605.09469v1 Announce Type: new Abstract: This paper explores the use of emojis in financial sentiment analysis, focusing on the social media platform StockTwits. Emojis, increasingly prevalent

researcharxiv-cs-cl
12 May 2026
Tutorials

First, Do No Harm: AI Supervisor Scaffolds Novice Growth in Counselor Education

DGX agent

arXiv:2508.09042v3 Announce Type: replace Abstract: The most dangerous mistakes a novice counselor makes are not the obvious ones: they are utterances that sound caring while quietly violating profess

tutorialsarxiv-cs-cl
12 May 2026
Agents

FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning

DGX agent

arXiv:2605.09932v1 Announce Type: new Abstract: Large language models can now process increasingly long inputs, yet their ability to effectively use information spread across long contexts remains lim

agentsarxiv-cs-cl
12 May 2026
Model Releases

Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling

DGX agent

arXiv:2512.02010v5 Announce Type: replace Abstract: As large language models have grown larger, interest has grown in low-precision numerical formats such as NVFP4 as a way to improve speed and reduce

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Frame In, Frame Out: Measuring Framing Bias in LLM-Generated News Summaries

DGX agent

arXiv:2505.05406v2 Announce Type: replace Abstract: News headlines and summaries shape how events are interpreted through selective emphasis and omission, a phenomenon commonly referred to as framing.

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM Collectives

DGX agent

arXiv:2605.09027v1 Announce Type: new Abstract: In multi-agent systems (MAS), a single deceptive agent can nullify all gains of an agentic AI collective and evade deployed defenses. However, existing

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

GIFT: Guided Importance-Aware Fine-Tuning for Diffusion Language Models

DGX agent

arXiv:2509.20863v3 Announce Type: replace Abstract: Diffusion models have recently shown strong potential in language modeling, offering faster generation compared to traditional autoregressive approa

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

GLiNER-Relex: A Unified Framework for Joint Named Entity Recognition and Relation Extraction

DGX agent

arXiv:2605.10108v1 Announce Type: new Abstract: Joint named entity recognition (NER) and relation extraction (RE) is a fundamental task in natural language processing for constructing knowledge graphs

model-releasesarxiv-cs-cl
12 May 2026
Agents

GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression

DGX agent

arXiv:2605.09100v1 Announce Type: new Abstract: Text embedding and generative tasks are usually trained separately based on large language models (LLMs) nowadays. This causes a large amount of trainin

agentsarxiv-cs-cl
12 May 2026
Model Releases

Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking

DGX agent

arXiv:2605.10893v1 Announce Type: new Abstract: Large vision-language models suffer from visual ungroundedness: they can produce a fluent, confident, and even correct response driven entirely by langu

model-releasesarxiv-cs-cl
12 May 2026
Research

Grounded Satirical Generation with RAG

DGX agent

arXiv:2605.10853v1 Announce Type: new Abstract: Humor generation remains challenging task for Large Language Models (LLMs), due to their subjective nature. We focus on satire, a form of humor strongly

researcharxiv-cs-cl
12 May 2026
Model Releases

Hint Tuning: Less Data Makes Better Reasoners

DGX agent

arXiv:2605.08665v1 Announce Type: new Abstract: Large reasoning models achieve high accuracy through extended chain-of-thought but generate 5--8 more tokens than necessary, applying verbose reasoning

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models

DGX agent

arXiv:2404.18923v5 Announce Type: replace Abstract: We introduce Holmes, a new benchmark designed to assess language models (LMs) linguistic competence - their unconscious understanding of linguistic

model-releasesarxiv-cs-cl
12 May 2026
Research

How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits

DGX agent

arXiv:2605.08348v1 Announce Type: new Abstract: The circuits framework in mechanistic interpretability aims to identify causally important sparse subgraphs of model components, typically evaluated by

researcharxiv-cs-cl
12 May 2026
Research

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue

DGX agent

arXiv:2605.10199v1 Announce Type: new Abstract: Full-duplex spoken dialogue requires a model to keep listening while generating its own spoken response. This is challenging for large language models (

researcharxiv-cs-cl
12 May 2026
Research

ICT-NLP at SemEval-2026 Task 3: Less Is More -- Multilingual Encoder with Joint Training and Adaptive Ensemble for Dimensional Aspect Sentiment Regression

DGX agent

arXiv:2605.10560v1 Announce Type: new Abstract: This paper describes our system to SemEval-2026 Task 3 Track A Subtask 1 on Dimensional Aspect Sentiment Regression (DimASR). We propose a lightweight a

researcharxiv-cs-cl
12 May 2026
Model Releases

Incremental Multilingual Text2Cypher with Adapter Combination

DGX agent

arXiv:2601.16097v2 Announce Type: replace Abstract: Large Language Models enable users to access database using natural language interfaces using tools like Text2SQL, Text2SPARQL, and Text2Cypher, whi

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables

DGX agent

arXiv:2605.10039v1 Announce Type: cross Abstract: Frontier coding agents read configuration files (CLAUDE.md, AGENTS.md, Cursor Rules) at session start and are expected to follow the conventions insid

model-releasesarxiv-cs-cl
12 May 2026
Safety

Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration

DGX agent

arXiv:2602.03677v2 Announce Type: replace Abstract: Modality following is the ability to selectively leverage multimodal contexts based on user instructions. It is fundamental to the safety and reliab

safetyarxiv-cs-cl
12 May 2026
Model Releases

jina-embeddings-v5-omni: Text-Geometry-Preserving Multimodal Embeddings via Frozen-Tower Composition

DGX agent

arXiv:2605.08384v1 Announce Type: new Abstract: In this work, we introduce frozen-encoder model composition, a novel approach to multimodal embedding models. We build on the VLM-style architecture, in

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

DGX agent

arXiv:2605.09635v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in K-12 education, yet existing benchmarks such as C-Eval, CMMLU, GaokaoBench, and EduEval mainly eva

model-releasesarxiv-cs-cl
12 May 2026
Research

Language-Conditioned Visual Grounding with CLIP Multilingual

DGX agent

arXiv:2605.09060v1 Announce Type: new Abstract: Multilingual vision-language models exhibit systematic performance gaps across languages, but the mechanism remains ambiguous: cross-language divergence

researcharxiv-cs-cl
12 May 2026
Model Releases

Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes

DGX agent

arXiv:2605.09751v1 Announce Type: new Abstract: Trainable input embedding tables are a standard component of modern language models. We ask whether they are actually necessary at the input interface.

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Large Language Models as Students Who Think Aloud: Overly Coherent, Verbose, and Confident

DGX agent

arXiv:2602.01015v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly embedded in AI-based tutoring systems. Can they faithfully model novice reasoning and metacognitive ju

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LEAF-SQL: Level-wise Exploration with Adaptive Fine-graining for Text-to-SQL Skeleton Prediction

DGX agent

arXiv:2605.09295v1 Announce Type: new Abstract: Text-to-SQL translates natural language questions into executable SQL queries, enabling intuitive database access for non-experts. While large language

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining

DGX agent

arXiv:2605.10504v1 Announce Type: new Abstract: A causal-decoder block is hierarchical: lower layers build the residual basis that upper layers attend over. We identify a failure mode in GPT pretraini

model-releasesarxiv-cs-cl
12 May 2026
Research

Learning More from Less: Exploiting Counterfactuals for Data-Efficient Chart Understanding

DGX agent

arXiv:2605.10855v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated remarkable progress in chart understanding, largely driven by supervised fine-tuning (SFT) on increasing

researcharxiv-cs-cl
12 May 2026
Safety

Learning to Stay Safe: Adaptive Regularization Against Safety Degradation during Fine-Tuning

DGX agent

arXiv:2602.17546v2 Announce Type: replace Abstract: Instruction-following language models are trained to be helpful and safe, yet their safety behavior can deteriorate under benign fine-tuning and wor

safetyarxiv-cs-cl
12 May 2026
Research

Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants

DGX agent

arXiv:2508.16070v3 Announce Type: replace Abstract: Approximately 283 million people worldwide live with visual impairments, motivating increasing research into leveraging Visual Language Models (VLMs

researcharxiv-cs-cl
12 May 2026
Safety

Let the Target Select for Itself: Data Selection via Target-Aligned Paths

DGX agent

arXiv:2605.09404v1 Announce Type: cross Abstract: Targeted data selection aims to identify training samples from a large candidate pool that improve performance on a specific downstream task. Many rec

safetyarxiv-cs-cl
12 May 2026
Model Releases

LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments

DGX agent

arXiv:2605.10779v1 Announce Type: cross Abstract: The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content s

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LLiMba: Sardinian on a Single GPU -- Adapting a 3B Language Model to a Vanishing Romance Language

DGX agent

arXiv:2605.09015v1 Announce Type: new Abstract: Sardinian, a Romance language with roughly one million speakers, has minimal presence in modern NLP. Commercial services do not support it, and current

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LLM Agents Already Know When to Call Tools -- Even Without Reasoning

DGX agent

arXiv:2605.09252v1 Announce Type: new Abstract: Tool-augmented LLM agents tend to call tools indiscriminately, even when the model can answer directly. Each unnecessary call wastes API fees and latenc

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LLMs with in-context learning for Algorithmic Theoretical Physics

DGX agent

arXiv:2605.08212v1 Announce Type: cross Abstract: There is an increasing number of algorithmic computations in theoretical physics. These, while conceptually simple, can nevertheless be time-consuming

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories

DGX agent

arXiv:2509.20909v2 Announce Type: replace Abstract: Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, where benchmark contamination can make memor

model-releasesarxiv-cs-cl
12 May 2026
← Previous
1…9899100101102…161
Next →