AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
10 Apr 2026

Linear Representations of Hierarchical Concepts in Language Models

TutorialsDGX agent

arXiv:2604.07886v1 Announce Type: new Abstract: We investigate how and to what extent hierarchical relations (e.g., Japan subset Eastern Asia subset Asia) are encoded in the internal representat

LLM-Based Data Generation and Clinical Skills Evaluation for Low-Resource French OSCEs

ApplicationsDGX agent

arXiv:2604.08126v1 Announce Type: new Abstract: Objective Structured Clinical Examinations (OSCEs) are the standard method for assessing medical students' clinical and communication skills through str

LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization

ResearchDGX agent

arXiv:2510.13907v3 Announce Type: replace Abstract: Large language models (LLMs) are highly sensitive to prompts, but most automatic prompt optimization (APO) methods assume access to ground-truth ref


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers

ResearchDGX agent

arXiv:2604.07822v1 Announce Type: new Abstract: We study implicit reasoning, i.e. the ability to combine knowledge or rules within a single forward pass. While transformer-based large language models

MARCH: Evaluating the Intersection of Ambiguity Interpretation and Multi-hop Inference

Model ReleasesDGX agent

arXiv:2509.22750v3 Announce Type: replace Abstract: Real-world multi-hop QA is naturally linked with ambiguity, where a single query can trigger multiple reasoning paths that require independent resol

MemReader: From Passive to Active Extraction for Long-Term Agent Memory

SafetyDGX agent

arXiv:2604.07877v1 Announce Type: new Abstract: Long-term memory is fundamental for personalized and autonomous agents, yet populating it remains a bottleneck. Existing systems treat memory extraction

Mina: A Multilingual LLM-Powered Legal Assistant Agent for Bangladesh for Empowering Access to Justice

AgentsDGX agent

arXiv:2511.08605v3 Announce Type: replace Abstract: Bangladesh's low-income population faces major barriers to affordable legal advice due to complex legal language, procedural opacity, and high costs

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale

Model ReleasesDGX agent

arXiv:2604.04771v2 Announce Type: replace-cross Abstract: Current document parsing methods advance primarily through model architecture innovation, while systematic engineering of training data remain

Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing

Model ReleasesDGX agent

arXiv:2604.07747v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve low-k reasoning accuracy while narrowing solution coverage on challenging math que

ModeX: Evaluator-Free Best-of-N Selection for Open-Ended Generation

Model ReleasesDGX agent

arXiv:2601.02535v2 Announce Type: replace Abstract: Selecting a single high-quality output from multiple stochastic generations remains a fundamental challenge for large language models (LLMs), partic

More Capable, Less Cooperative? When LLMs Fail At Zero-Cost Collaboration

AgentsDGX agent

arXiv:2604.07821v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly coordinate in multi-agent systems, yet we lack an understanding of where and why cooperation failures m

OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence

ApplicationsDGX agent

arXiv:2604.07296v2 Announce Type: replace Abstract: Spatial understanding is a fundamental cornerstone of human-level intelligence. Nonetheless, current research predominantly focuses on domain-specif

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks

SafetyDGX agent

arXiv:2604.08539v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) has emerged as the de facto Reinforcement Learning (RL) objective driving recent advancements in Multimodal

Optimal Decay Spectra for Linear Recurrences

ResearchDGX agent

arXiv:2604.07658v1 Announce Type: cross Abstract: Linear recurrent models offer linear-time sequence processing but often suffer from suboptimal long-range memory. We trace this to the decay spectrum:

ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents

AgentsDGX agent

arXiv:2604.07789v1 Announce Type: cross Abstract: Recent advances in language model (LM) agents have significantly improved automated software engineering (SWE). Prior work has proposed various agenti

OrgForge: A Multi-Agent Simulation Framework for Verifiable Synthetic Corporate Corpora

AgentsDGX agent

arXiv:2603.14997v2 Announce Type: replace Abstract: Building and evaluating enterprise AI systems requires synthetic organizational corpora that are internally consistent, temporally structured, and c

Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speech

ResearchDGX agent

arXiv:2512.24517v2 Announce Type: replace Abstract: Automatic speech transcripts are often delivered as unstructured word streams that impede readability and repurposing. We recast paragraph segmentat

PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory

Model ReleasesDGX agent

arXiv:2604.08000v1 Announce Type: cross Abstract: Proactivity is a core expectation for AGI. Prior work remains largely confined to laboratory settings, leaving a clear gap in real-world proactive age

PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning

SafetyDGX agent

arXiv:2508.09521v2 Announce Type: replace Abstract: Emotional support conversations require more than fluent responses. Supporters need to understand the seeker's situation and emotions, adopt an appr

PeReGrINE: Evaluating Personalized Review Fidelity with User Item Graph Context

Model ReleasesDGX agent

arXiv:2604.07788v1 Announce Type: cross Abstract: We introduce PeReGrINE, a benchmark and evaluation framework for personalized review generation grounded in graph-structured user--item evidence. PeRe

PIArena: A Platform for Prompt Injection Evaluation

ApplicationsDGX agent

arXiv:2604.08499v1 Announce Type: cross Abstract: Prompt injection attacks pose serious security risks across a wide range of real-world applications. While receiving increasing attention, the communi

PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch

Model ReleasesDGX agent

arXiv:2510.06670v2 Announce Type: replace Abstract: High-quality instruction data is critical for LLM alignment, yet existing open-source datasets often lack efficiency, requiring hundreds of thousand

Prompt reinforcing for long-term planning of large language models

Model ReleasesDGX agent

arXiv:2510.05921v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved remarkable success in a wide range of natural language processing tasks and can be adapted through prompt

Prune-Quantize-Distill: An Ordered Pipeline for Efficient Neural Network Compression

Model ReleasesDGX agent

arXiv:2604.04988v1 Announce Type: cross Abstract: Modern deployment often requires trading accuracy for efficiency under tight CPU and memory constraints, yet common compression proxies such as parame

Quantum Vision Theory Applied to Audio Classification for Deepfake Speech Detection

ResearchDGX agent

arXiv:2604.08104v1 Announce Type: new Abstract: We propose Quantum Vision (QV) theory as a new perspective for deep learning-based audio classification, applied to deepfake speech detection. Inspired

Rag Performance Prediction for Question Answering

ResearchDGX agent

arXiv:2604.07985v1 Announce Type: new Abstract: We address the task of predicting the gain of using RAG (retrieval augmented generation) for question answering with respect to not using it. We study t

Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs

ApplicationsDGX agent

arXiv:2604.07562v1 Announce Type: new Abstract: Unsupervised methods are widely used to induce latent semantic structure from large text collections, yet their outputs often contain incoherent, redund

Reasoning Graphs: Deterministic Agent Accuracy through Evidence-Centric Chain-of-Thought Feedback

AgentsDGX agent

arXiv:2604.07595v1 Announce Type: cross Abstract: Language model agents reason from scratch on every query: each time an agent retrieves evidence and deliberates, the chain of thought is discarded and

Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space

SafetyDGX agent

arXiv:2512.12623v3 Announce Type: replace-cross Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced cross-modal understanding and reasoning by incorpo

ReCellTy: Domain-Specific Knowledge Graph Retrieval-Augmented LLMs Reasoning Workflow for Single-Cell Annotation

ResearchDGX agent

arXiv:2505.00017v2 Announce Type: replace Abstract: With the rapid development of large language models (LLMs), their application to cell type annotation has drawn increasing attention. However, gener

ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework

SafetyDGX agent

arXiv:2604.07506v1 Announce Type: cross Abstract: Reward Models (RMs) are critical components in the Reinforcement Learning from Human Feedback (RLHF) pipeline, directly determining the alignment qual

Rethinking Data Mixing from the Perspective of Large Language Models

ResearchDGX agent

arXiv:2604.07963v1 Announce Type: new Abstract: Data mixing strategy is essential for large language model (LLM) training. Empirical evidence shows that inappropriate strategies can significantly redu

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs

Model ReleasesDGX agent

arXiv:2604.08003v1 Announce Type: cross Abstract: Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a dominant paradigm. Although recent LLM-based ASR models

SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive Thinking

ResearchDGX agent

arXiv:2604.07922v1 Announce Type: cross Abstract: Large Reasoning Models (LRMs) have revolutionized complex problem-solving, yet they exhibit a pervasive 'overthinking', generating unnecessarily long

sciwrite-lint: Verification Infrastructure for the Age of Science Vibe-Writing

Model ReleasesDGX agent

arXiv:2604.08501v1 Announce Type: cross Abstract: Science currently offers two options for quality assurance, both inadequate. Journal gatekeeping claims to verify both integrity and contribution, but

SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models

Model ReleasesDGX agent

arXiv:2506.01062v4 Announce Type: replace Abstract: We introduce SealQA, a new challenge benchmark for evaluating SEarch-Augmented Language models on fact-seeking questions where web search yields con

Search-R3: Unifying Reasoning and Embedding in Large Language Models

TutorialsDGX agent

arXiv:2510.07048v2 Announce Type: replace Abstract: Despite their remarkable natural language understanding capabilities, Large Language Models (LLMs) have been underutilized for retrieval tasks. We p

See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video LLMs

ResearchDGX agent

arXiv:2604.05650v2 Announce Type: replace Abstract: Video Large Language Models (Video-LLMs) excel in video understanding but suffer from high inference latency during autoregressive generation. Specu

Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts

SafetyDGX agent

arXiv:2604.08541v1 Announce Type: cross Abstract: Multimodal Mixture-of-Experts (MoE) models have achieved remarkable performance on vision-language tasks. However, we identify a puzzling phenomenon t

Seeing Like an AI: How LLMs Apply (and Misapply) Wikipedia Neutrality Norms

SafetyDGX agent

arXiv:2407.04183v4 Announce Type: replace Abstract: Large language models (LLMs) are trained on broad corpora and then used in communities with specialized norms. Is providing LLMs with community rule

SeLaR: Selective Latent Reasoning in Large Language Models

ResearchDGX agent

arXiv:2604.08299v1 Announce Type: new Abstract: Chain-of-Thought (CoT) has become a cornerstone of reasoning in large language models, yet its effectiveness is constrained by the limited expressivenes

Self-Debias: Self-correcting for Debiasing Large Language Models

SafetyDGX agent

arXiv:2604.08243v1 Announce Type: new Abstract: Although Large Language Models (LLMs) demonstrate remarkable reasoning capabilities, inherent social biases often cascade throughout the Chain-of-Though

Sell More, Play Less: Benchmarking LLM Realistic Selling Skill

Model ReleasesDGX agent

arXiv:2604.07054v2 Announce Type: replace Abstract: Sales dialogues require multi-turn, goal-directed persuasion under asymmetric incentives, which makes them a challenging setting for large language

Sensitivity-Positional Co-Localization in GQA Transformers

Model ReleasesDGX agent

arXiv:2604.07766v1 Announce Type: new Abstract: We investigate a fundamental structural question in Grouped Query Attention (GQA) transformers: do the layers most sensitive to task correctness coincid

SepSeq: A Training-Free Framework for Long Numerical Sequence Processing in LLMs

Local AiDGX agent

arXiv:2604.07737v1 Announce Type: new Abstract: While transformer-based Large Language Models (LLMs) theoretically support massive context windows, they suffer from severe performance degradation when

SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

AgentsDGX agent

arXiv:2604.08377v1 Announce Type: cross Abstract: Large language model (LLM) agents such as OpenClaw rely on reusable skills to perform complex tasks, yet these skills remain largely static after depl

Small Vision-Language Models are Smart Compressors for Long Video Understanding

Model ReleasesDGX agent

arXiv:2604.08120v1 Announce Type: cross Abstract: Adapting Multimodal Large Language Models (MLLMs) for hour-long videos is bottlenecked by context limits. Dense visual streams saturate token budgets

SOLAR: Communication-Efficient Model Adaptation via Subspace-Oriented Latent Adapter Reparametrization

Model ReleasesDGX agent

arXiv:2604.08368v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning (PEFT) methods, such as LoRA, enable scalable adaptation of foundation models by injecting low-rank adapters. However,

Splits! Flexible Sociocultural Linguistic Investigation at Scale

ResearchDGX agent

arXiv:2504.04640v3 Announce Type: replace Abstract: Variation in language use, shaped by speakers' sociocultural background and specific context of use, offers a rich lens into cultural perspectives,

Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution

ResearchDGX agent

arXiv:2604.07725v1 Announce Type: cross Abstract: We show that verifier-free evolution is bottlenecked by both diversity and efficiency: without external correction, repeated evolution accelerates col

Stacked from One: Multi-Scale Self-Injection for Context Window Extension

Model ReleasesDGX agent

arXiv:2603.04759v2 Announce Type: replace Abstract: The limited context window of contemporary large language models (LLMs) remains a primary bottleneck for their broader application across diverse do

Stay Focused: Problem Drift in Multi-Agent Debate

AgentsDGX agent

arXiv:2502.19559v3 Announce Type: replace Abstract: Multi-agent debate - multiple instances of large language models discussing problems in turn-based interaction - has shown promise for solving knowl

Stop Listening to Me! How Multi-turn Conversations Can Degrade LLM Diagnostic Reasoning

ApplicationsDGX agent

arXiv:2603.11394v2 Announce Type: replace Abstract: Patients and clinicians are increasingly using chatbots powered by large language models (LLMs) for healthcare inquiries. While state-of-the-art LLM

SubSearch: Intermediate Rewards for Unsupervised Guided Reasoning in Complex Retrieval

AgentsDGX agent

arXiv:2604.07415v1 Announce Type: cross Abstract: Large language models (LLMs) are probabilistic in nature and perform more reliably when augmented with external information. As complex queries often

Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding

Model ReleasesDGX agent

arXiv:2604.07753v1 Announce Type: cross Abstract: Empowering Large Multimodal Models (LMMs) with image generation often leads to catastrophic forgetting in understanding tasks due to severe gradient c

SYN-DIGITS: A Synthetic Control Framework for Calibrated Digital Twin Simulation

SafetyDGX agent

arXiv:2604.07513v1 Announce Type: cross Abstract: AI-based persona simulation -- often referred to as digital twin simulation -- is increasingly used for market research, recommender systems, and soci

Synthetic Data for any Differentiable Target

SafetyDGX agent

arXiv:2604.08423v1 Announce Type: new Abstract: What are the limits of controlling language models via synthetic training data? We develop a reinforcement learning (RL) primitive, the Dataset Policy G

TEC: A Collection of Human Trial-and-error Trajectories for Problem Solving

TutorialsDGX agent

arXiv:2604.06734v2 Announce Type: replace Abstract: Trial-and-error is a fundamental strategy for humans to solve complex problems and a necessary capability for Artificial Intelligence (AI) systems o

TEMPER: Testing Emotional Perturbation in Quantitative Reasoning

Model ReleasesDGX agent

arXiv:2604.07801v1 Announce Type: new Abstract: Large language models are trained and evaluated on quantitative reasoning tasks written in clean, emotionally neutral language. However, real-world quer

Testimole-Conversational: A 30-Billion-Word Italian Discussion Board Corpus (1996-2024) for Language Modeling and Sociolinguistic Research

ResearchDGX agent

arXiv:2602.14819v2 Announce Type: replace Abstract: We present 'Testimole-conversational' a massive collection of discussion boards messages in the Italian language. The large size of the corpus, more

← Previous
1…125126127128
Next →