AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
5 May 2026

CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features

Model ReleasesDGX agent

arXiv:2508.12535v3 Announce Type: replace Abstract: Sparse Autoencoders (SAEs) can extract interpretable features from large language models (LLMs) without supervision. However, their effectiveness in

CoSpaDi: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning

Model ReleasesDGX agent

arXiv:2509.22075v5 Announce Type: replace Abstract: Post-training compression of large language models (LLMs) often relies on low-rank weight approximations that represent each column of the weight ma

Counting as a minimal probe of language model reliability

ResearchDGX agent

arXiv:2605.02028v1 Announce Type: new Abstract: Large language models perform strongly on benchmarks in mathematical reasoning, coding and document analysis, suggesting a broad ability to follow instr


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

CP-SynC: Multi-Agent Zero-Shot Constraint Modeling in MiniZinc with Synthesized Checkers

Model ReleasesDGX agent

arXiv:2605.01675v1 Announce Type: cross Abstract: Constraint Programming (CP) is a powerful paradigm for solving combinatorial problems, yet translating natural language problem descriptions into exec

Creating and Evaluating Figurative Language Dataset for Sindhi

Model ReleasesDGX agent

arXiv:2605.01323v1 Announce Type: new Abstract: In this article, we introduce SiNFluD, a novel benchmark dataset for Sindhi figurative language classification. We first collect raw text from various b

CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation

Model ReleasesDGX agent

arXiv:2603.01865v3 Announce Type: replace Abstract: LLM-as-judge evaluation has become standard practice for open-ended model assessment; however, judges exhibit systematic biases that cannot be avera

Decoding-Time Debiasing via Process Reward Models: From Controlled Fill-in to Open-Ended Generation

Model ReleasesDGX agent

arXiv:2605.02348v1 Announce Type: new Abstract: Large language models pick up social biases from the data they are trained on and carry those biases into downstream applications, often reinforcing ste

DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning

HardwareDGX agent

arXiv:2510.09883v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieve state-of-the-art performance on challenging benchmarks by generating long chains of intermediate steps, but th

Democratizing the medieval English legal tradition

Model ReleasesDGX agent

arXiv:2605.00977v1 Announce Type: cross Abstract: The record of the beginning of the most widespread legal system in the world is contained in millions of pages of handwritten text. Most of the record

Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages

ResearchDGX agent

arXiv:2605.02608v1 Announce Type: new Abstract: Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-

DIAGRAMS: A Review Framework for Reasoning-Level Attribution in Diagram QA

ResearchDGX agent

arXiv:2605.00905v1 Announce Type: new Abstract: Diagram question answering (Diagram QA) requires reasoning-level attribution that links each question-answer pair to all visual regions needed to derive

Do Large Language Models Plan Answer Positions? Position Bias in Multiple-Choice Question Generation

SafetyDGX agent

arXiv:2605.01846v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate multiple-choice questions (MCQs), where correct answers should ideally be uniformly distr

EditPropBench: Measuring Factual Edit Propagation in Scientific Manuscripts

Model ReleasesDGX agent

arXiv:2605.02083v1 Announce Type: new Abstract: Local factual edits in scientific manuscripts often create non-local revision obligations. If a dataset changes from 215 to 80 documents, claims such as

Efficient Reasoning with Hidden Thinking

Model ReleasesDGX agent

arXiv:2501.19201v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) reasoning has become a powerful framework for improving complex problem-solving capabilities in Multimodal Large Language Mod

EGAD: Entropy-Guided Adaptive Distillation for Token-Level Knowledge Transfer

TutorialsDGX agent

arXiv:2605.01732v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable performance across diverse domains, yet their enormous computational and memory requirements hinde

Embedding-based In-Context Prompt Training for Enhancing LLMs as Text Encoders

Model ReleasesDGX agent

arXiv:2605.01372v1 Announce Type: new Abstract: Large language models (LLMs) have been widely explored for embedding generation. While recent studies show that in-context learning (ICL) effectively en

Energy-Based Constraint Networks: Learning Structural Coherence Across Modalities

Local AiDGX agent

arXiv:2605.00960v1 Announce Type: cross Abstract: We introduce energy-based constraint networks -- a modality-agnostic architecture that learns structural coherence from contrastive pairs. The system

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.02073v1 Announce Type: new Abstract: Mathematical reasoning is a key benchmark for large language models. Reinforcement learning is a standard post-training mechanism for improving the reas

Enhancing Game Review Sentiment Classification on Steam Platform with Attention-Based BiLSTM

ResearchDGX agent

arXiv:2605.01315v1 Announce Type: new Abstract: This paper investigates sentiment classification of Steam game reviews using an attention-based Bidirectional Long Short-Term Memory (BiLSTM) model. Usi

Enhancing Judgment Document Generation via Agentic Legal Information Collection and Rubric-Guided Optimization

Model ReleasesDGX agent

arXiv:2605.02011v1 Announce Type: new Abstract: Automating the drafting of judgment documents is pivotal to judicial efficiency, yet it remains challenging due to the dual requirements of comprehensiv

Extracting memorized pieces of (copyrighted) books from open-weight language models

Model ReleasesDGX agent

arXiv:2505.12546v5 Announce Type: replace Abstract: Plaintiffs and defendants in copyright lawsuits over generative AI often make sweeping, opposing claims about the extent to which large language mod

Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture

Model ReleasesDGX agent

arXiv:2605.01567v1 Announce Type: cross Abstract: Large language model (LLM) coding agents increasingly operate over repositories, terminals, tests, and execution traces across long software-engineeri

Fight Poison with Poison: Enhancing Robustness in Few-shot Machine-Generated Text Detection with Adversarial Training

ResearchDGX agent

arXiv:2605.02374v1 Announce Type: cross Abstract: Machine-generated text (MGT) detection is critical for regulating online information ecosystems, yet existing detectors often underperform in few-shot

Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2508.15202v2 Announce Type: replace Abstract: Process Reward Models (PRMs) supervise intermediate reasoning steps in large language models (LLMs), but existing PRMs are mainly trained on general

Fine-Tuning Pre-Trained Code Models for AI-Generated Code Detection

ResearchDGX agent

arXiv:2605.01596v1 Announce Type: new Abstract: This paper describes the system submitted by team extbf{Archaeology} to SemEval-2026 Task~13 on AI-generated code detection. The shared task consists of

Flexi-LoRA with Input-Adaptive Ranks: Efficient Finetuning for Speech and Reasoning Tasks

Model ReleasesDGX agent

arXiv:2605.01959v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning methods like Low-Rank Adaptation (LoRA) have become essential for deploying large language models, yet their static pa

FlexSQL: Flexible Exploration and Execution Make Better Text-to-SQL Agents

Model ReleasesDGX agent

arXiv:2605.02815v1 Announce Type: new Abstract: Text-to-SQL over large analytical databases requires navigating complex schemas, resolving ambiguous queries, and grounding decisions in actual data. Mo

Focus on the Core: Empowering Diffusion Large Language Models by Self-Contrast

Model ReleasesDGX agent

arXiv:2605.01373v1 Announce Type: new Abstract: The iterative denoising paradigm of Diffusion Large Language Models (DLMs) endows them with a distinct advantage in global context modeling. However, cu

Foundation Models to Unlock Real-World Evidence from Nationwide Medical Claims

SafetyDGX agent

arXiv:2605.02740v1 Announce Type: cross Abstract: Evidence derived from large-scale real-world data (RWD) is increasingly informing regulatory evaluation and healthcare decision-making. Administrative

FT-RAG: A Fine-grained Retrieval-Augmented Generation Framework for Complex Table Reasoning

Model ReleasesDGX agent

arXiv:2605.01495v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by grounding responses in external knowledge during inference. However, conve

FunFuzz: An LLM-Powered Evolutionary Fuzzing Framework

ResearchDGX agent

arXiv:2605.02789v1 Announce Type: cross Abstract: Modern fuzzers increasingly use Large Language Models (LLMs) to generate structured inputs, but LLM-driven fuzzing is sensitive to prompt initializati

Fuzzy Fingerprinting Encoder Pre-trained Language Models for Emotion Recognition in Conversations: Human Assessment and Validity Study

ResearchDGX agent

arXiv:2605.02665v1 Announce Type: new Abstract: In Emotion Recognition in Conversations (ERC), model decisions should align with nuanced human perception and ideally provide insights on the classifica

Generative Interfaces for Language Models

ResearchDGX agent

arXiv:2508.19227v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly seen as assistants, copilots, and consultants, capable of supporting a wide range of tasks through nat

GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language Models

TutorialsDGX agent

arXiv:2605.01256v1 Announce Type: new Abstract: A promising paradigm for adapting instruction-tuned language models is to learn task-specific updates on a pretrained base model and subsequently merge

GR-Ben: A General Reasoning Benchmark for Evaluating Process Reward Models

Model ReleasesDGX agent

arXiv:2605.01203v1 Announce Type: cross Abstract: Currently, process reward models (PRMs) have exhibited remarkable potential for test-time scaling. Since large language models (LLMs) regularly genera

GRAIL: A Deep-Granularity Hybrid Resonance Framework for Real-Time Agent Discovery via SLM-Enhanced Indexing

AgentsDGX agent

arXiv:2605.02489v1 Announce Type: cross Abstract: As the ecosystem of Large Language Model (LLM)-based agents expands rapidly, efficient and accurate Agent Discovery becomes a critical bottleneck for

Graph Query Generation with Constraint-guided Large Language Agents

ApplicationsDGX agent

arXiv:2605.00845v1 Announce Type: cross Abstract: Knowledge Graph Question Answering (KGQA) has advanced through structured query generation, yet most efforts target RDF/SPARQL, leaving Cypher and pro

GRAVITY: Architecture-Agnostic Structured Anchoring for Long-Horizon Conversational Memory

ResearchDGX agent

arXiv:2605.01688v1 Announce Type: new Abstract: Long-horizon conversational agents rely on memory systems with increasingly sophisticated retrieval mechanisms. However, retrieved fragments are typical

Growing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate

Model ReleasesDGX agent

arXiv:2507.07129v3 Announce Type: replace-cross Abstract: We study a constrained training regime for decoder-only Transformers in which the token interface is fixed, previously trained dense blocks ar

H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models

ApplicationsDGX agent

arXiv:2605.00847v1 Announce Type: new Abstract: Representing and navigating hierarchy is a fundamental primitive of reasoning. Large language models have demonstrated proficiency in a wide variety of

Hallucination Detection in LLMs with Topological Divergence on Attention Graphs

ResearchDGX agent

arXiv:2504.10063v4 Announce Type: replace Abstract: Hallucination, i.e., generating factually incorrect content, remains a critical challenge for large language models (LLMs). We introduce TOHA, a TOp

Hallucinations Undermine Trust; Metacognition is a Way Forward

AgentsDGX agent

arXiv:2605.01428v1 Announce Type: new Abstract: Despite significant strides in factual reliability, errors -- often termed hallucinations -- remain a major concern for generative AI, especially as LLM

HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs

Model ReleasesDGX agent

arXiv:2605.02443v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse natural language processing tasks, yet they remain susceptible to

HeteroRAG: A Heterogeneous Retrieval-Augmented Generation Framework for Medical Vision Language Tasks

SafetyDGX agent

arXiv:2508.12778v2 Announce Type: replace Abstract: Medical large vision-language Models (Med-LVLMs) have shown promise in clinical applications but suffer from factual inaccuracies and unreliable out

Hey, That's My Data! Token-Only Dataset Inference in Large Language Models

ResearchDGX agent

arXiv:2506.06057v2 Announce Type: replace Abstract: Large Language Models (LLMs) rely on massive training datasets, often including proprietary data, which raises concerns about unauthorized usage and

How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP

ResearchDGX agent

arXiv:2411.05527v3 Announce Type: replace Abstract: Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such

How Prompts Move Language Model Behavior: Frames, Salience, and Construal as Semantic Control

ResearchDGX agent

arXiv:2512.12688v3 Announce Type: replace-cross Abstract: Prompt engineering is widely used to shape large language model behavior, yet it is often treated as a practical heuristic rather than as a fo

How Well Can We Decode Vowels from Auditory EEG -- A Rigorous Cross-Subject Benchmark with Honest Assessment

Model ReleasesDGX agent

arXiv:2605.00865v1 Announce Type: cross Abstract: EEG based phoneme decoding is promising for brain computer interfaces, but many prior studies rely on within subject evaluation, small cohorts, or wea

How^{2}: How to learn from procedural How-to questions

AgentsDGX agent

arXiv:2510.11144v2 Announce Type: replace-cross Abstract: An agent facing a planning problem can use answers to how-to questions to reduce uncertainty and fill knowledge gaps, helping it solve both cu

Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs

Model ReleasesDGX agent

arXiv:2502.16435v4 Announce Type: replace-cross Abstract: Humans develop perception through a bottom-up hierarchy: from basic primitives and Gestalt principles to high-level semantics. In contrast, cu

Implicature in Interaction: Understanding Implicature Improves Alignment in Human-LLM Interaction

SafetyDGX agent

arXiv:2510.25426v2 Announce Type: replace Abstract: The rapid advancement of Large Language Models (LLMs) is positioning language at the core of human-computer interaction (HCI). We argue that advanci

Improving Factuality in LLMs via Inference-Time Knowledge Graph Construction

ResearchDGX agent

arXiv:2509.03540v3 Announce Type: replace Abstract: Large Language Models (LLMs) often struggle with producing factually consistent answers due to limitations in their parametric memory. Retrieval-Aug

InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition

ResearchDGX agent

arXiv:2605.02364v1 Announce Type: new Abstract: Upweighting high-quality data in LLM pretraining often improves performance, but in datalimited regimes, especially under overtraining, stronger upweigh

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression

SafetyDGX agent

arXiv:2605.01402v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) struggle with numerical regression under long-tailed target distributions. Token-level supervised fine-tuning (

Interpretable Difficulty-Aware Knowledge Tracing in Tutor-Student Dialogues

ResearchDGX agent

arXiv:2605.01097v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have led to the development of AI-powered tutoring systems that provide interactive support via dialogue

IPS: In-Prompt Process Supervision for Short Video Content Moderation

SafetyDGX agent

arXiv:2412.15251v3 Announce Type: replace Abstract: Multimodal large language models (MLLMs) are effective at capturing the semantics of short video content; however, they often fail to attend to the

Is It Novel and Why? Fine-Grained Patent Novelty Prediction Based on Passage Retrieval

ResearchDGX agent

arXiv:2605.02392v1 Announce Type: new Abstract: Novelty assessment is a critical yet complex task in the examination process for patent acceptance, requiring examiners to determine whether an inventio

jina-vlm: Small Multilingual Vision Language Model

Model ReleasesDGX agent

arXiv:2512.04032v3 Announce Type: replace Abstract: We present jina-vlm, a token-efficient 2.4B parameter vision-language model that achieves state-of-the-art multilingual VQA performance among open 2

Latent Trajectory Dynamics in Large Language Models: A Manifold Evolution Framework with Empirical Validation

Model ReleasesDGX agent

arXiv:2505.20340v3 Announce Type: replace Abstract: Understanding how latent representations evolve during generation is a central open problem in large language model interpretability. We introduce e

LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference

HardwareDGX agent

arXiv:2605.01058v1 Announce Type: cross Abstract: Layer-aligned distillation and convergence-based early exit represent two predominant computational efficiency paradigms for transformer inference; ye

← Previous
1…8788899091…129
Next →