AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Model Releases

RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025)

DGX agent

arXiv:2604.16392v1 Announce Type: cross Abstract: AI in Education research increasingly relies on authentic, curriculum-grounded assessment data, yet large, well-structured exam corpora remain scarce

model-releasesarxiv-cs-cl
21 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

RoTRAG: Rule of Thumb Reasoning for Conversation Harm Detection with Retrieval-Augmented Generation

DGX agent

arXiv:2604.17301v1 Announce Type: new Abstract: Detecting harmful content in multi turn dialogue requires reasoning over the full conversational context rather than isolated utterances. However, most

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

S-GRPO: Unified Post-Training for Large Vision-Language Models

DGX agent

arXiv:2604.16557v1 Announce Type: cross Abstract: Current post-training methodologies for adapting Large Vision-Language Models (LVLMs) generally fall into two paradigms: Supervised Fine-Tuning (SFT)

safetyarxiv-cs-cl
21 Apr 2026
Safety

SaFeR-Steer: Evolving Multi-Turn MLLMs via Synthetic Bootstrapping and Feedback Dynamics

DGX agent

arXiv:2604.16358v1 Announce Type: cross Abstract: MLLMs are increasingly deployed in multi-turn settings, where attackers can escalate unsafe intent through the evolving visual-text history and exploi

safetyarxiv-cs-cl
21 Apr 2026
Safety

Safety, Security, and Cognitive Risks in State-Space Models: A Systematic Threat Analysis with Spectral, Stateful, and Capacity Attacks

DGX agent

arXiv:2604.16424v1 Announce Type: cross Abstract: State-Space Models (SSMs) -- structured SSMs (S4, S4D, DSS, S5), selective SSMs (Mamba, Mamba-2), and hybrid architectures (Jamba) -- are deployed in

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection

DGX agent

arXiv:2601.05403v2 Announce Type: replace Abstract: Large language models (LLMs) have been widely applied across various domains of finance. Since their training data are largely derived from human-au

model-releasesarxiv-cs-cl
21 Apr 2026
Agents

Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding

DGX agent

arXiv:2510.15253v3 Announce Type: replace Abstract: Document understanding is critical for applications from financial analysis to scientific discovery. Current approaches, whether OCR-based pipelines

agentsarxiv-cs-cl
21 Apr 2026
Model Releases

Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent Collaboration

DGX agent

arXiv:2505.21471v2 Announce Type: replace Abstract: With the rapid advancement of post-training techniques for reasoning and information seeking, large language models (LLMs) can incorporate a large q

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Scaling Test-Time Compute for Agentic Coding

DGX agent

arXiv:2604.16529v1 Announce Type: cross Abstract: Test-time scaling has become a powerful way to improve large language models. However, existing methods are best suited to short, bounded outputs that

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows

DGX agent

arXiv:2505.19897v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have extended their impact beyond Natural Language Processing, substantially fostering the development of interdi

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

SciImpact: A Multi-Dimensional, Multi-Field Benchmark for Scientific Impact Prediction

DGX agent

arXiv:2604.17141v1 Announce Type: new Abstract: The rapid growth of scientific literature calls for automated methods to assess and predict research impact. Prior work has largely focused on citation-

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Screen Before You Interpret: A Portable Validity Protocol for Benchmark-Based LLM Confidence Signals

DGX agent

arXiv:2604.17714v1 Announce Type: new Abstract: LLM confidence signals are used for abstention, routing, and safety-critical decisions. No standard practice exists for checking whether a confidence si

model-releasesarxiv-cs-cl
21 Apr 2026
Research

Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework

DGX agent

arXiv:2602.19549v2 Announce Type: replace Abstract: Visual Document Retrieval (VDR), which aims to retrieve relevant pages within vast corpora of visually-rich documents, is of significance in current

researcharxiv-cs-cl
21 Apr 2026
Agents

Seeing Isn't Believing: Mitigating Belief Inertia via Active Intervention in Embodied Agents

DGX agent

arXiv:2604.17252v1 Announce Type: new Abstract: Recent advancements in large language models (LLMs) have enabled agents to tackle complex embodied tasks through environmental interaction. However, the

agentsarxiv-cs-cl
21 Apr 2026
Research

Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning

DGX agent

arXiv:2604.17433v1 Announce Type: new Abstract: Self-consistency (SC) is a popular technique for improving the reasoning accuracy of large language models by aggregating multiple sampled outputs, but

researcharxiv-cs-cl
21 Apr 2026
Local Ai

Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement

DGX agent

arXiv:2411.15115v3 Announce Type: replace-cross Abstract: Recent text-to-video (T2V) diffusion models have made remarkable progress in generating high-quality videos. However, they often struggle to a

local-aiarxiv-cs-cl
21 Apr 2026
Research

Semantic Density Effect (SDE): Maximizing Information Per Token Improves LLM Accuracy

DGX agent

arXiv:2604.17659v1 Announce Type: new Abstract: We introduce the Semantic Density Effect (SDE): the empirical finding that prompts carrying higher semantic information per token consistently produce m

researcharxiv-cs-cl
21 Apr 2026
Research

Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code Reasoning

DGX agent

arXiv:2505.13353v4 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed for understanding large codebases, but whether they understand operational semantics of long

researcharxiv-cs-cl
21 Apr 2026
Agents

SERM: Self-Evolving Relevance Model with Agent-Driven Learning from Massive Query Streams

DGX agent

arXiv:2601.09515v2 Announce Type: replace Abstract: Due to the dynamically evolving nature of real-world query streams, relevance models struggle to generalize to practical search scenarios. A sophist

agentsarxiv-cs-cl
21 Apr 2026
Research

Sessa: Selective State Space Attention

DGX agent

arXiv:2604.18580v1 Announce Type: cross Abstract: Modern sequence models are dominated by Transformers, where self-attention mixes information from the visible context in an input-dependent way. Howev

researcharxiv-cs-cl
21 Apr 2026
Applications

SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe

DGX agent

arXiv:2410.05248v4 Announce Type: replace Abstract: To acquire instruction-following capabilities, large language models (LLMs) undergo instruction tuning, where they are trained on instruction-respon

applicationsarxiv-cs-cl
21 Apr 2026
Local Ai

SignDPO: Multi-level Direct Preference Optimisation for Skeleton-based Gloss-free Sign Language Translation

DGX agent

arXiv:2604.18034v1 Announce Type: new Abstract: We present SignDPO, a novel multi-level Direct Preference Optimisation (DPO) framework designed to enhance the alignment of skeleton-based Sign Language

local-aiarxiv-cs-cl
21 Apr 2026
Agents

SkillX: Automatically Constructing Skill Knowledge Bases for Agents

DGX agent

arXiv:2604.04804v2 Announce Type: replace Abstract: Learning from experience is critical for building capable large language model (LLM) agents, yet prevailing self-evolving paradigms remain inefficie

agentsarxiv-cs-cl
21 Apr 2026
Research

SmoGVLM: A Small, Graph-enhanced Vision-Language Model

DGX agent

arXiv:2604.16517v1 Announce Type: cross Abstract: Large vision-language models (VLMs) achieve strong performance on multimodal tasks but often suffer from hallucination and poor grounding in knowledge

researcharxiv-cs-cl
21 Apr 2026
Research

Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models

DGX agent

arXiv:2506.18141v3 Announce Type: replace Abstract: We identify semantically coherent, context-consistent network components in large language models (LLMs) using coactivation of sparse autoencoder (S

researcharxiv-cs-cl
21 Apr 2026
Model Releases

SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?

DGX agent

arXiv:2601.04029v2 Announce Type: replace Abstract: Large Audio-Language Models (LALMs) as judges have emerged as a prominent approach for evaluating speech generation quality, yet their ability to as

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Spec-o3: A Tool-Augmented Vision-Language Agent for Rare Celestial Object Candidate Vetting via Automated Spectral Inspection

DGX agent

arXiv:2601.06498v2 Announce Type: replace Abstract: Due to the limited generalization and interpretability of deep learning classifiers, The final vetting of rare celestial object candidates still rel

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding

DGX agent

arXiv:2509.24328v2 Announce Type: replace Abstract: LLMs have low GPU efficiency and high latency due to autoregressive decoding. Speculative decoding (SD) mitigates this using a small draft model to

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation

DGX agent

arXiv:2601.04638v2 Announce Type: replace Abstract: Medical consultations are intrinsically speech-centric. However, most prior works focus on long-text-based interactions, which are cumbersome and pa

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks

DGX agent

arXiv:2604.17771v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance on natural language to SQL (NL2SQL) benchmarks, yet their reported accuracy may be inflate

model-releasesarxiv-cs-cl
21 Apr 2026
Research

SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation

DGX agent

arXiv:2512.21204v2 Announce Type: replace Abstract: Human infants, with only a few hundred hours of speech exposure, acquire basic units of new languages, highlighting a striking efficiency gap compar

researcharxiv-cs-cl
21 Apr 2026
Safety

SpiralThinker: Latent Reasoning through an Iterative Process with Text-Latent Interleaving

DGX agent

arXiv:2511.08983v2 Announce Type: replace Abstract: Recent advances in large reasoning models have been driven by reinforcement learning and test-time scaling, accompanied by growing interest in laten

safetyarxiv-cs-cl
21 Apr 2026
Research

Spotlights and Blindspots: Evaluation Machine-Generated Text Detection

DGX agent

arXiv:2604.16607v1 Announce Type: new Abstract: With the rise of generative language models, machine-generated text detection has become a critical challenge. A wide variety of models is available, bu

researcharxiv-cs-cl
21 Apr 2026
Safety

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models

DGX agent

arXiv:2604.16995v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a promising paradigm for training reasoning-oriented models by leveraging rule-based reward signals. However,

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

SQL Query Engine: A Self-Healing LLM Pipeline for Natural Language to PostgreSQL Translation

DGX agent

arXiv:2604.16511v1 Announce Type: cross Abstract: We present SQL Query Engine, an open-source, self-hosted service that translates natural language questions into validated PostgreSQL queries through

model-releasesarxiv-cs-cl
21 Apr 2026
Research

Stability-Weighted Decoding for Diffusion Language Models

DGX agent

arXiv:2604.17068v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) enable parallel text generation by iteratively denoising a fully masked sequence, unmasking a subset of masked t

researcharxiv-cs-cl
21 Apr 2026
Research

Stable Language Guidance for Vision-Language-Action Models

DGX agent

arXiv:2601.04052v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have demonstrated impressive capabilities in generalized robotic control; however, they remain notoriously

researcharxiv-cs-cl
21 Apr 2026
Model Releases

STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs

DGX agent

arXiv:2604.18177v1 Announce Type: new Abstract: Benchmarks are often used as a standard to understand LLM capabilities in different domains. However, aggregate benchmark scores provide limited insight

model-releasesarxiv-cs-cl
21 Apr 2026
Research

StageMem: Lifecycle-Managed Memory for Language Models

DGX agent

arXiv:2604.16774v1 Announce Type: new Abstract: Long-horizon language model systems increasingly rely on persistent memory, yet many current designs still treat memory primarily as a static store: wri

researcharxiv-cs-cl
21 Apr 2026
Safety

StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation

DGX agent

arXiv:2601.04740v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied in specialized domains such as finance and healthcare, where they introduce unique safety risk

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning

DGX agent

arXiv:2604.18401v1 Announce Type: new Abstract: General agents have given rise to phenomenal applications such as OpenClaw and Claude Code. As these agent systems (a.k.a. Harnesses) strive for bolder

model-releasesarxiv-cs-cl
21 Apr 2026
Applications

Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions

DGX agent

arXiv:2604.17358v1 Announce Type: new Abstract: While recent Spoken Language Models (SLMs) have been actively deployed in real-world scenarios, they lack the capability to discern Third-Party Interrup

applicationsarxiv-cs-cl
21 Apr 2026
Research

Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs

DGX agent

arXiv:2602.11528v2 Announce Type: replace-cross Abstract: Recent studies have shown that large language models (LLMs) can infer private user attributes (e.g., age, location, gender) from user-generate

researcharxiv-cs-cl
21 Apr 2026
Safety

Structure-Aware Diversity Pursuit as an AI Safety Strategy against Homogenization

DGX agent

arXiv:2601.06116v2 Announce Type: replace-cross Abstract: Generative AI models reproduce the biases in the training data and can further amplify them through mode collapse. We refer to the resulting h

safetyarxiv-cs-cl
21 Apr 2026
Research

Style over Story: Measuring LLM Narrative Preferences via Structured Selection

DGX agent

arXiv:2510.02025v4 Announce Type: replace Abstract: We introduce a constraint-selection-based experiment design for measuring narrative preferences of Large Language Models (LLMs). This design offers

researcharxiv-cs-cl
21 Apr 2026
Safety

SynopticBench: Evaluating Vision-Language Models on Generating Weather Forecast Discussions of the Future

DGX agent

arXiv:2604.16451v1 Announce Type: new Abstract: Recent advances in visual-language models (VLMs) have led to significant improvements in a plethora of complex multimodal tasks like image captioning, r

safetyarxiv-cs-cl
21 Apr 2026
Research

Synthetic Data Generation for Training Diversified Commonsense Reasoning Models

DGX agent

arXiv:2603.18361v2 Announce Type: replace Abstract: Conversational agents are required to respond to their users not only with high quality (i.e. commonsense bearing) responses, but also considering m

researcharxiv-cs-cl
21 Apr 2026
Safety

Synthia: Scalable Grounded Persona Generation from Social Media Data

DGX agent

arXiv:2507.14922v2 Announce Type: replace Abstract: Persona-driven simulations are increasingly used in computational social science, yet their validity critically depends on the fidelity of the under

safetyarxiv-cs-cl
21 Apr 2026
← Previous
1…137138139140141…160
Next →