AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
21 Apr 2026

ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning

Model ReleasesDGX agent

arXiv:2603.05863v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have revolutionized code generation, standard ``System 1'' approaches that generate solutions in a single forward

Reinforced Efficient Reasoning via Semantically Diverse Exploration

Model ReleasesDGX agent

arXiv:2601.05053v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has proven effective in enhancing the reasoning of large language models (LLMs). Monte C

Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2601.02970v2 Announce Type: replace Abstract: Self-Consistency improves reasoning reliability through multi-sample aggregation, but incurs substantial inference cost. Adaptive self-consistency m

Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning

SafetyDGX agent

arXiv:2601.14750v3 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting has achieved remarkable success in unlocking the reasoning capabilities of Large Language Models (LLMs). Although C

Representation-Guided Parameter-Efficient LLM Unlearning

Model ReleasesDGX agent

arXiv:2604.17396v1 Announce Type: new Abstract: Large Language Models (LLMs) often memorize sensitive or harmful information, necessitating effective machine unlearning techniques. While existing para

RePrompT: Recurrent Prompt Tuning for Integrating Structured EHR Encoders with Large Language Models

TutorialsDGX agent

arXiv:2604.17725v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong promise for mining Electronic Health Records (EHRs) by reasoning over longitudinal clinical information t

ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Model ReleasesDGX agent

arXiv:2503.21248v3 Announce Type: replace Abstract: Large language models (LLMs) have shown potential in assisting scientific research, yet their ability to discover high-quality research hypotheses r

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring

SafetyDGX agent

arXiv:2512.12069v3 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) are vulnerable to a growing array of multimodal jailbreak attacks, necessitating defenses that are both g

Rethinking Meeting Effectiveness: A Benchmark and Framework for Temporal Fine-grained Automatic Meeting Effectiveness Evaluation

Model ReleasesDGX agent

arXiv:2604.17260v1 Announce Type: new Abstract: Evaluating meeting effectiveness is crucial for improving organizational productivity. Current approaches rely on post-hoc surveys that yield a single c

ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering

Model ReleasesDGX agent

arXiv:2510.09351v2 Announce Type: replace Abstract: While Small Language Models (SLMs) have demonstrated promising performance on an increasingly wide array of commonsense reasoning benchmarks, curren

Retrieval-Augmented Multimodal Model for Fake News Detection

SafetyDGX agent

arXiv:2604.18112v1 Announce Type: new Abstract: In recent years, multimodal multidomain fake news detection has garnered increasing attention. Nevertheless, this direction presents two significant cha

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF

SafetyDGX agent

arXiv:2604.17769v1 Announce Type: new Abstract: Ensuring the safety of large language models (LLMs) requires robust red teaming, yet the systematic synthesis of high-quality toxic data remains under-e

Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models

Model ReleasesDGX agent

arXiv:2604.16593v1 Announce Type: new Abstract: We present SemanticQA, an evaluation suite designed to assess language models (LMs) in semantic phrase processing tasks. The benchmark consolidates exis

Revisiting Entropy in Reinforcement Learning for Large Reasoning Models

Local AiDGX agent

arXiv:2511.05993v3 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a prominent paradigm for enhancing the reasoning capabilities of large language

REZE: Representation Regularization for Domain-adaptive Text Embedding Pre-finetuning

SafetyDGX agent

arXiv:2604.17257v1 Announce Type: new Abstract: Recent text embedding models are often adapted to specialized domains via contrastive pre-finetuning (PFT) on a naive collection of scattered, heterogen

River-LLM: Large Language Model Seamless Exit Based on KV Share

TutorialsDGX agent

arXiv:2604.18396v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated exceptional performance across diverse domains but are increasingly constrained by high inference latency

Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models

Model ReleasesDGX agent

arXiv:2602.14466v2 Announce Type: replace Abstract: With natural language generation becoming a popular use case for language models, the Bias Benchmark for Question-Answering (BBQ) has grown to be an

RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian

Model ReleasesDGX agent

arXiv:2604.17134v1 Announce Type: new Abstract: We present RoIt-XMASA, a multilingual dataset that extends the Cross-lingual Multi-domain Amazon Sentiment Analysis to Italian and Romanian, comprising

RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025)

Model ReleasesDGX agent

arXiv:2604.16392v1 Announce Type: cross Abstract: AI in Education research increasingly relies on authentic, curriculum-grounded assessment data, yet large, well-structured exam corpora remain scarce

RoTRAG: Rule of Thumb Reasoning for Conversation Harm Detection with Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2604.17301v1 Announce Type: new Abstract: Detecting harmful content in multi turn dialogue requires reasoning over the full conversational context rather than isolated utterances. However, most

S-GRPO: Unified Post-Training for Large Vision-Language Models

SafetyDGX agent

arXiv:2604.16557v1 Announce Type: cross Abstract: Current post-training methodologies for adapting Large Vision-Language Models (LVLMs) generally fall into two paradigms: Supervised Fine-Tuning (SFT)

SaFeR-Steer: Evolving Multi-Turn MLLMs via Synthetic Bootstrapping and Feedback Dynamics

SafetyDGX agent

arXiv:2604.16358v1 Announce Type: cross Abstract: MLLMs are increasingly deployed in multi-turn settings, where attackers can escalate unsafe intent through the evolving visual-text history and exploi

Safety, Security, and Cognitive Risks in State-Space Models: A Systematic Threat Analysis with Spectral, Stateful, and Capacity Attacks

SafetyDGX agent

arXiv:2604.16424v1 Announce Type: cross Abstract: State-Space Models (SSMs) -- structured SSMs (S4, S4D, DSS, S5), selective SSMs (Mamba, Mamba-2), and hybrid architectures (Jamba) -- are deployed in

Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection

Model ReleasesDGX agent

arXiv:2601.05403v2 Announce Type: replace Abstract: Large language models (LLMs) have been widely applied across various domains of finance. Since their training data are largely derived from human-au

Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding

AgentsDGX agent

arXiv:2510.15253v3 Announce Type: replace Abstract: Document understanding is critical for applications from financial analysis to scientific discovery. Current approaches, whether OCR-based pipelines

Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent Collaboration

Model ReleasesDGX agent

arXiv:2505.21471v2 Announce Type: replace Abstract: With the rapid advancement of post-training techniques for reasoning and information seeking, large language models (LLMs) can incorporate a large q

Scaling Test-Time Compute for Agentic Coding

Model ReleasesDGX agent

arXiv:2604.16529v1 Announce Type: cross Abstract: Test-time scaling has become a powerful way to improve large language models. However, existing methods are best suited to short, bounded outputs that

ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows

Model ReleasesDGX agent

arXiv:2505.19897v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have extended their impact beyond Natural Language Processing, substantially fostering the development of interdi

SciImpact: A Multi-Dimensional, Multi-Field Benchmark for Scientific Impact Prediction

Model ReleasesDGX agent

arXiv:2604.17141v1 Announce Type: new Abstract: The rapid growth of scientific literature calls for automated methods to assess and predict research impact. Prior work has largely focused on citation-

Screen Before You Interpret: A Portable Validity Protocol for Benchmark-Based LLM Confidence Signals

Model ReleasesDGX agent

arXiv:2604.17714v1 Announce Type: new Abstract: LLM confidence signals are used for abstention, routing, and safety-critical decisions. No standard practice exists for checking whether a confidence si

Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework

ResearchDGX agent

arXiv:2602.19549v2 Announce Type: replace Abstract: Visual Document Retrieval (VDR), which aims to retrieve relevant pages within vast corpora of visually-rich documents, is of significance in current

Seeing Isn't Believing: Mitigating Belief Inertia via Active Intervention in Embodied Agents

AgentsDGX agent

arXiv:2604.17252v1 Announce Type: new Abstract: Recent advancements in large language models (LLMs) have enabled agents to tackle complex embodied tasks through environmental interaction. However, the

Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning

ResearchDGX agent

arXiv:2604.17433v1 Announce Type: new Abstract: Self-consistency (SC) is a popular technique for improving the reasoning accuracy of large language models by aggregating multiple sampled outputs, but

Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement

Local AiDGX agent

arXiv:2411.15115v3 Announce Type: replace-cross Abstract: Recent text-to-video (T2V) diffusion models have made remarkable progress in generating high-quality videos. However, they often struggle to a

Semantic Density Effect (SDE): Maximizing Information Per Token Improves LLM Accuracy

ResearchDGX agent

arXiv:2604.17659v1 Announce Type: new Abstract: We introduce the Semantic Density Effect (SDE): the empirical finding that prompts carrying higher semantic information per token consistently produce m

Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code Reasoning

ResearchDGX agent

arXiv:2505.13353v4 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed for understanding large codebases, but whether they understand operational semantics of long

SERM: Self-Evolving Relevance Model with Agent-Driven Learning from Massive Query Streams

AgentsDGX agent

arXiv:2601.09515v2 Announce Type: replace Abstract: Due to the dynamically evolving nature of real-world query streams, relevance models struggle to generalize to practical search scenarios. A sophist

Sessa: Selective State Space Attention

ResearchDGX agent

arXiv:2604.18580v1 Announce Type: cross Abstract: Modern sequence models are dominated by Transformers, where self-attention mixes information from the visible context in an input-dependent way. Howev

SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe

ApplicationsDGX agent

arXiv:2410.05248v4 Announce Type: replace Abstract: To acquire instruction-following capabilities, large language models (LLMs) undergo instruction tuning, where they are trained on instruction-respon

SignDPO: Multi-level Direct Preference Optimisation for Skeleton-based Gloss-free Sign Language Translation

Local AiDGX agent

arXiv:2604.18034v1 Announce Type: new Abstract: We present SignDPO, a novel multi-level Direct Preference Optimisation (DPO) framework designed to enhance the alignment of skeleton-based Sign Language

SkillX: Automatically Constructing Skill Knowledge Bases for Agents

AgentsDGX agent

arXiv:2604.04804v2 Announce Type: replace Abstract: Learning from experience is critical for building capable large language model (LLM) agents, yet prevailing self-evolving paradigms remain inefficie

SmoGVLM: A Small, Graph-enhanced Vision-Language Model

ResearchDGX agent

arXiv:2604.16517v1 Announce Type: cross Abstract: Large vision-language models (VLMs) achieve strong performance on multimodal tasks but often suffer from hallucination and poor grounding in knowledge

Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models

ResearchDGX agent

arXiv:2506.18141v3 Announce Type: replace Abstract: We identify semantically coherent, context-consistent network components in large language models (LLMs) using coactivation of sparse autoencoder (S

SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?

Model ReleasesDGX agent

arXiv:2601.04029v2 Announce Type: replace Abstract: Large Audio-Language Models (LALMs) as judges have emerged as a prominent approach for evaluating speech generation quality, yet their ability to as

Spec-o3: A Tool-Augmented Vision-Language Agent for Rare Celestial Object Candidate Vetting via Automated Spectral Inspection

Model ReleasesDGX agent

arXiv:2601.06498v2 Announce Type: replace Abstract: Due to the limited generalization and interpretability of deep learning classifiers, The final vetting of rare celestial object candidates still rel

Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding

SafetyDGX agent

arXiv:2509.24328v2 Announce Type: replace Abstract: LLMs have low GPU efficiency and high latency due to autoregressive decoding. Speculative decoding (SD) mitigates this using a small draft model to

SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation

Model ReleasesDGX agent

arXiv:2601.04638v2 Announce Type: replace Abstract: Medical consultations are intrinsically speech-centric. However, most prior works focus on long-text-based interactions, which are cumbersome and pa

SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks

Model ReleasesDGX agent

arXiv:2604.17771v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance on natural language to SQL (NL2SQL) benchmarks, yet their reported accuracy may be inflate

SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation

ResearchDGX agent

arXiv:2512.21204v2 Announce Type: replace Abstract: Human infants, with only a few hundred hours of speech exposure, acquire basic units of new languages, highlighting a striking efficiency gap compar

SpiralThinker: Latent Reasoning through an Iterative Process with Text-Latent Interleaving

SafetyDGX agent

arXiv:2511.08983v2 Announce Type: replace Abstract: Recent advances in large reasoning models have been driven by reinforcement learning and test-time scaling, accompanied by growing interest in laten

Spotlights and Blindspots: Evaluation Machine-Generated Text Detection

ResearchDGX agent

arXiv:2604.16607v1 Announce Type: new Abstract: With the rise of generative language models, machine-generated text detection has become a critical challenge. A wide variety of models is available, bu

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models

SafetyDGX agent

arXiv:2604.16995v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a promising paradigm for training reasoning-oriented models by leveraging rule-based reward signals. However,

SQL Query Engine: A Self-Healing LLM Pipeline for Natural Language to PostgreSQL Translation

Model ReleasesDGX agent

arXiv:2604.16511v1 Announce Type: cross Abstract: We present SQL Query Engine, an open-source, self-hosted service that translates natural language questions into validated PostgreSQL queries through

Stability-Weighted Decoding for Diffusion Language Models

ResearchDGX agent

arXiv:2604.17068v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) enable parallel text generation by iteratively denoising a fully masked sequence, unmasking a subset of masked t

Stable Language Guidance for Vision-Language-Action Models

ResearchDGX agent

arXiv:2601.04052v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have demonstrated impressive capabilities in generalized robotic control; however, they remain notoriously

STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs

Model ReleasesDGX agent

arXiv:2604.18177v1 Announce Type: new Abstract: Benchmarks are often used as a standard to understand LLM capabilities in different domains. However, aggregate benchmark scores provide limited insight

StageMem: Lifecycle-Managed Memory for Language Models

ResearchDGX agent

arXiv:2604.16774v1 Announce Type: new Abstract: Long-horizon language model systems increasingly rely on persistent memory, yet many current designs still treat memory primarily as a static store: wri

StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation

SafetyDGX agent

arXiv:2601.04740v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied in specialized domains such as finance and healthcare, where they introduce unique safety risk

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.18401v1 Announce Type: new Abstract: General agents have given rise to phenomenal applications such as OpenClaw and Claude Code. As these agent systems (a.k.a. Harnesses) strive for bolder

Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions

ApplicationsDGX agent

arXiv:2604.17358v1 Announce Type: new Abstract: While recent Spoken Language Models (SLMs) have been actively deployed in real-world scenarios, they lack the capability to discern Third-Party Interrup

← Previous
1…110111112113114…129
Next →