AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
2 Jun 2026

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding

SafetyDGX agent

arXiv:2606.00564v1 Announce Type: cross Abstract: While on-policy distillation offers dense supervision for training small reasoning models, its optimization dynamics in the multimodal domain remain u

Deep networks learn to parse uniform-depth context-free languages from local statistics

TutorialsDGX agent

arXiv:2602.06065v3 Announce Type: replace-cross Abstract: Understanding how the structure of language can be learned from sentences alone is a central question in both cognitive science and machine le

Deep Research as Rubric for Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.01091v1 Announce Type: new Abstract: Open-ended reasoning and long-form generation tasks lack reliable automatic verification signals for reward-based policy optimization. Rubrics offer a p


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

DeSQ: Decomposition-based SPARQL Query Generation

ResearchDGX agent

arXiv:2606.00203v1 Announce Type: new Abstract: Dominant approaches to Knowledge Base Question Answering (KBQA) fall into two categories. First is the generation of a formal query that suffers from br

DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding

ResearchDGX agent

arXiv:2606.02091v1 Announce Type: new Abstract: Block diffusion speculative decoding accelerates LLM inference by predicting all tokens within a block simultaneously for the target model to verify in

Digging Up Citations: FOSSIL, a Dataset and Workflow for Reference Extraction in Law and the Humanities

ResearchDGX agent

arXiv:2606.01109v1 Announce Type: cross Abstract: Citation extraction tools are designed for the structured end-of-document bibliographies of the natural sciences, but law and humanities scholarship c

Disentangling Similarity and Relatedness in Topic Models

Model ReleasesDGX agent

arXiv:2603.10619v2 Announce Type: replace Abstract: The recent success of large pre-trained language models (PLMs) has motivated their integration into topic modeling. However, PLM-augmented topic mod

Do Gender Cues Affect LLM Value Trade-offs? Evidence from a Controlled Decision Benchmark

Model ReleasesDGX agent

arXiv:2606.02214v1 Announce Type: new Abstract: Large language models are increasingly used in value-sensitive decision settings, where irrelevant demographic cues should not alter judgments. We const

Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs

Model ReleasesDGX agent

arXiv:2606.00477v1 Announce Type: new Abstract: Unified multimodal models (UMMs) have emerged as a promising paradigm for general-purpose multimodal intelligence. As they are deployed in real-world ap

Don't Read Everything: A Curvature-Conditioned Query for Linear Attention

Local AiDGX agent

arXiv:2606.01294v1 Announce Type: new Abstract: Linear attention reduces the quadratic cost of softmax attention by maintaining a recurrent fast-weight state, but it consistently lags on in-context re

DrugClaw and DrugAudit: A Primary-Source-Grounded Agent and Authority-Aware Benchmark for Drug-Information Question Answering

Model ReleasesDGX agent

arXiv:2606.01434v1 Announce Type: new Abstract: Drug-information question answering is a high-stakes setting where hallucinated facts can mislead clinical decision-making and the provenance of each ci

Efficient RAG with Intent-Aware Retrieval and Semantics-Preserving Chunking

Model ReleasesDGX agent

arXiv:2606.01240v1 Announce Type: new Abstract: The demand for powerful instruction following and reasoning capability of large language models (LLMs) has promoted rapid development of retrieval-augme

Empathy Applicability Modeling for General Health Queries

Model ReleasesDGX agent

arXiv:2601.09696v2 Announce Type: replace Abstract: LLMs are increasingly being integrated into clinical workflows, yet they often lack clinical empathy, an essential aspect of effective doctor-patien

Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification

ResearchDGX agent

arXiv:2606.01679v1 Announce Type: new Abstract: Multimodal LLMs are increasingly used to assist scientific peer review, where a core requirement is verifying whether claims in a paper are supported by

Escaping the Mode Lottery: Multi-Response Training Improves Language Model Generalization

Model ReleasesDGX agent

arXiv:2606.00544v1 Announce Type: cross Abstract: Modern language-model fine-tuning typically pairs each prompt with a single response, even though many prompts admit multiple valid completions. This

Evaluating the Reversal Curse in Model Editing

Model ReleasesDGX agent

arXiv:2310.10322v3 Announce Type: replace Abstract: Large language models (LLMs) are prone to hallucinate unintended text due to false or outdated knowledge. Since retraining LLMs is resource intensiv

ExpWeaver: LLM Agents Learn from Experience via Latent RAG

Model ReleasesDGX agent

arXiv:2606.01041v1 Announce Type: new Abstract: Experience learning has achieved promising results in enhancing LLM agent planning and reasoning by integrating past interactions as reusable knowledge.

Eyettention II: A Dual-Sequence Architecture for Modeling Fixation Location, Within-Word Landing Position, and Fixation Duration in Reading

HardwareDGX agent

arXiv:2606.01964v1 Announce Type: new Abstract: The way our eyes move while reading provides valuable insights into both the reader's cognitive processes and the properties of the text. In particular,

FigSIM: A Dataset for Fine-grained Suicide Severity and Figurative Language in Suicide Memes

Model ReleasesDGX agent

arXiv:2606.02523v1 Announce Type: new Abstract: Suicide memes are memes used to express suicide-related thoughts or comment on suicide-related issues. Suicide memes are increasingly common on social m

Finding What Matters: Anchoring Context Knowledge with Evolving Indices for Iterative Retrieval

TutorialsDGX agent

arXiv:2601.16462v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) has become a dominant paradigm for mitigating hallucinations in Large Language Models (LLMs) by incorporating e

Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models

ApplicationsDGX agent

arXiv:2602.23197v2 Announce Type: replace Abstract: Transformer-based large language models exhibit in-context learning, enabling adaptation to downstream tasks via few-shot prompting with demonstrati

FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search

Model ReleasesDGX agent

arXiv:2606.00660v1 Announce Type: new Abstract: Agentic search requires language model agents to explore many sources and answer complex information-seeking questions. Scaling test-time compute is a p

ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution

ResearchDGX agent

arXiv:2602.03203v2 Announce Type: replace Abstract: Recently, large language models (LLMs) have shown remarkable reasoning abilities by producing long reasoning traces. However, as the sequence length

French parsing enhanced with a word clustering method based on a syntactic lexicon

ResearchDGX agent

arXiv:2606.00634v1 Announce Type: new Abstract: This article evaluates the integration of data extracted from a French syntactic lexicon, the Lexicon-Grammar (Gross, 1994), into a probabilistic parser

From Empathy to Personalized Empathy: Adapting Empathetic Strategies to Individual Users

ResearchDGX agent

arXiv:2606.00728v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed in long-term interactions with users, empathy has become an increasingly important capability.

From Global to Local: Learning Context-Aware Graph Representations for Document Classification and Summarization

ResearchDGX agent

arXiv:2603.00021v2 Announce Type: replace Abstract: Recent NLP systems commonly represent documents as linear token sequences. Although this captures sequential order, it can hinder modeling long-rang

From Outliers to Errors: Auditing Pali-to-English LLM Translations with Multi-Reference Adjudication

Model ReleasesDGX agent

arXiv:2606.01136v1 Announce Type: new Abstract: Single-score translation metrics can conflate legitimate variation with error, a problem especially acute for classical languages where multiple defensi

From Unfamiliar to Familiar: Detecting Pre-training Data via Gradient Deviations in Large Language Models

Model ReleasesDGX agent

arXiv:2603.04828v2 Announce Type: replace Abstract: Pre-training data detection for LLMs is essential for addressing copyright concerns and mitigating benchmark contamination. Existing methods mainly

GateKD: Confidence-Gated Closed-Loop Distillation for Robust Reasoning

ResearchDGX agent

arXiv:2605.13136v2 Announce Type: replace Abstract: Distilling multi-step reasoning abilities from large language models (LLMs) into compact student models remains challenging due to noisy rationales,

GeistBERT: Breathing Life into German NLP

ResearchDGX agent

arXiv:2506.11903v5 Announce Type: replace Abstract: Advances in transformer-based language models have highlighted the benefits of language-specific pre-training on high-quality corpora. In this conte

Geometric Latent Reasoning Induces Shorter Generations in LLMs

ResearchDGX agent

arXiv:2606.02248v1 Announce Type: new Abstract: Large language models solve complex problems by generating lengthy chains of explicit reasoning tokens. While effective, this makes reasoning expensive,

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

Model ReleasesDGX agent

arXiv:2510.24081v2 Announce Type: replace Abstract: To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and

GottBERT: a pure German Language Model

ResearchDGX agent

arXiv:2012.02110v2 Announce Type: replace Abstract: Pre-trained language models have significantly advanced natural language processing (NLP), especially with the introduction of BERT and its optimize

Graph-Augmented Retrieval for Cross-Entity Financial Sentiment Analysis: A Comparative Study

ApplicationsDGX agent

arXiv:2606.00062v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become foundational for grounding large language models in domain-specific corpora, yet conventional vector-bas

Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language Translation

SafetyDGX agent

arXiv:2510.18439v3 Announce Type: replace Abstract: Hallucination, where models generate fluent text unsupported by visual evidence, remains a major flaw in vision-language models and is particularly

HalleluBERT: Let Every Token That Has Meaning Bear Its Weight

Model ReleasesDGX agent

arXiv:2510.21372v2 Announce Type: replace Abstract: Transformer-based models have advanced NLP, yet Hebrew still lacks a RoBERTa encoder that is trained at scale and released in both base and large va

HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems

SafetyDGX agent

arXiv:2606.01779v1 Announce Type: new Abstract: LLM agents are increasingly expected to operate across heterogeneous task regimes that require distinct execution paradigms. This challenges fixed agent

HERO'S JOURNEY: Testing Complex Rule Induction with Text Games

Model ReleasesDGX agent

arXiv:2606.02556v1 Announce Type: new Abstract: We introduce HERO'S JOURNEY, a benchmark for rule induction in goal-directed episodic tasks, where agents must infer hidden rules from demonstrations an

HMPO: Hybrid Median-length Policy Optimization for Chain-of-Thought Compression

SafetyDGX agent

arXiv:2606.01934v1 Announce Type: cross Abstract: Large language models achieve remarkable performance via extended chain-of-thought (CoT) reasoning, yet this lengthy process incurs substantial infere

How AI Fails: An Interactive Pedagogical Tool for Demonstrating Dialectal Bias in Automated Toxicity Models

Model ReleasesDGX agent

arXiv:2511.06676v3 Announce Type: replace Abstract: Now that AI-driven moderation has become pervasive in everyday life, we often hear claims that 'the AI is biased'. While this is often said jokingly

How Far Do Auto-Interpretation Labels Generalize: A Controlled Study Across Languages, Scripts, and Rewordings

ResearchDGX agent

arXiv:2606.00356v1 Announce Type: new Abstract: Sparse autoencoder (SAE) features are increasingly used to interpret language models, with auto-generated natural-language labels serving as the primary

How to Correctly Report LLM-as-a-Judge Evaluations

Model ReleasesDGX agent

arXiv:2511.21140v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are widely used as scalable evaluators of model responses in lieu of human annotators. However, imperfect sensiti

HypothesisMed: Inference-Time Answer Fusion and Structured Hypothesis-Space Reporting for Biomedical Question Answering

Model ReleasesDGX agent

arXiv:2606.00971v1 Announce Type: new Abstract: Biomedical question answering with large language models is commonly evaluated using answer accuracy, but answer accuracy alone does not indicate whethe

'I Strongly Suspect This Website Is a Scam': Benchmarking PII Leakage and Detection without Defense in Autonomous Web Agents

Model ReleasesDGX agent

arXiv:2606.00497v1 Announce Type: cross Abstract: Deceptive web content, widely instantiated across the internet and commonly known as extit{social-engineering attacks}, manipulates autonomous web age

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications

Model ReleasesDGX agent

arXiv:2606.00750v1 Announce Type: new Abstract: Recent advances in visual language models have enabled autonomous agents for complex reasoning, tool use, and document understanding. However, existing

IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs

ResearchDGX agent

arXiv:2606.00875v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for tasks involving creative problem solving and idea generation. However, there is a lack of consens

InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models

ApplicationsDGX agent

arXiv:2606.02161v1 Announce Type: cross Abstract: Video Large Language Models (Video-LLMs) achieve strong performance in video understanding, but their excessive visual tokens bring substantial comput

Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning

SafetyDGX agent

arXiv:2606.00755v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards improves the reasoning ability of large language models, but often suffers from entropy collapse, in whic

Interpreto: An Explainability Library for Transformers

ResearchDGX agent

arXiv:2512.09730v3 Announce Type: replace Abstract: Interpreto is an open-source Python library for interpreting HuggingFace language models, from early BERT variants to LLMs. It provides two compleme

Investigating and Alleviating Harm Amplification in LLM Interactions

Model ReleasesDGX agent

arXiv:2606.02423v1 Announce Type: new Abstract: Large language models (LLMs) can serve as helpful assistants, yet they can equally function as harm amplifiers that enable malicious users to achieve ha

K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

Model ReleasesDGX agent

arXiv:2606.02404v1 Announce Type: new Abstract: Frontier model evaluations are shifting from foundational capabilities (e.g., instruction following and reasoning) toward compositional, agentic ones, b

KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual Matrices

Model ReleasesDGX agent

arXiv:2601.21579v2 Announce Type: replace Abstract: The success of Hyper-Connections (HC) in neural networks (NN) has also highlighted issues related to training instability and restricted scalability

LaSR: Context-Aware Speech Recognition via Latent Reasoning

Model ReleasesDGX agent

arXiv:2606.00507v1 Announce Type: new Abstract: Recent advances in Speech Large Language Models (Speech LLMs) have significantly enhanced spoken language understanding and reasoning. However, their co

Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning

TutorialsDGX agent

arXiv:2511.07910v2 Announce Type: replace Abstract: Large Language Models (LLMs) achieve excellent performance in natural language reasoning tasks through pre-training on vast unstructured text, enabl

Learning from Saturated Data: Signals Beyond Correctness for LLM Training

ResearchDGX agent

arXiv:2606.01436v1 Announce Type: new Abstract: The growing capabilities of large language models (LLMs) have led to the saturation of many benchmarks and training datasets used to improve them. Motiv

Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation

AgentsDGX agent

arXiv:2602.03619v2 Announce Type: replace Abstract: Nowadays, developing reliable DeepResearch-style long-form report generation remains challenging, as training and evaluation lack verifiable reward

Learning to Retrieve: Dual-Level Long-Term Memory for Text-to-SQL Agents

Model ReleasesDGX agent

arXiv:2606.00547v1 Announce Type: new Abstract: Interactive text-to-SQL agents solve database tasks through multi-turn interactions involving schema exploration, query execution, feedback interpretati

Lessons from the Trenches on Reproducible Evaluation of Language Models

ResearchDGX agent

arXiv:2405.14782v3 Announce Type: replace Abstract: Reliable evaluation of language models (LMs) remains an open challenge. Re- searchers and engineers face methodological issues such as the sensitivi

LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding

ResearchDGX agent

arXiv:2602.23881v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to propose candidate t

LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation

SafetyDGX agent

arXiv:2603.09403v2 Announce Type: replace Abstract: Validating evaluation metrics for NLG typically relies on expensive and time-consuming human annotations, which predominantly exist only for English

← Previous
1…4849505152…129
Next →