AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
30 Apr 2026

Rethinking the Harmonic Loss via Non-Euclidean Distance Layers

ResearchDGX agent

arXiv:2603.10225v3 Announce Type: replace-cross Abstract: Cross-entropy loss has long been the standard choice for training deep neural networks, yet it suffers from interpretability limitations, unbo

Retrieval-Augmented LLMs for Evidence Localization in Clinical Trial Recruitment from Longitudinal EHR Narratives

Model ReleasesDGX agent

arXiv:2604.05190v2 Announce Type: replace-cross Abstract: Screening patients for enrollment is a well-known, labor-intensive bottleneck that leads to under-enrollment and, ultimately, trial failures.

RetroMotion: Retrocausal Motion Forecasting Models are Instructable

AgentsDGX agent

arXiv:2505.20414v2 Announce Type: replace-cross Abstract: Motion forecasts of road users (i.e., agents) vary in complexity depending on the number of agents, scene constraints, and interactions. In pa


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

reward-lens: A Mechanistic Interpretability Library for Reward Models

Model ReleasesDGX agent

arXiv:2604.26130v1 Announce Type: cross Abstract: Every RLHF-trained language model is shaped by a reward model, yet the mechanistic interpretability toolkit -- logit lens, direct logit attribution, a

Risk Reporting for Developers' Internal AI Model Use

SafetyDGX agent

arXiv:2604.24966v1 Announce Type: cross Abstract: Frontier AI companies first deploy their most advanced models internally, for weeks or months of safety testing, evaluation, and iteration, before a p

Robust Federated Learning under Adversarial Attacks via Loss-Based Client Clustering

ResearchDGX agent

arXiv:2508.12672v4 Announce Type: replace-cross Abstract: Federated Learning (FL) enables collaborative model training across multiple clients without sharing private data. We consider FL scenarios wh

Rule-based High-Level Coaching for Goal-Conditioned Reinforcement Learning in Search-and-Rescue UAV Missions Under Limited-Simulation Training

SafetyDGX agent

arXiv:2604.26833v1 Announce Type: cross Abstract: This paper presents a hierarchical decision-making framework for unmanned aerial vehicle (UAV) missions motivated by search-and-rescue (SAR) scenarios

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment

Model ReleasesDGX agent

arXiv:2601.04389v2 Announce Type: replace-cross Abstract: Current safety evaluations of large language models (LLMs) create a dangerous illusion of universal protection by aggregating harms under gene

SciHorizon-DataEVA: An Agentic System for AI-Readiness Evaluation of Heterogeneous Scientific Data

AgentsDGX agent

arXiv:2604.26645v1 Announce Type: new Abstract: AI-for-Science (AI4Science) is increasingly transforming scientific discovery by embedding machine learning models into prediction, simulation, and hypo

SciMDR: Advancing Scientific Multimodal Document Reasoning

Model ReleasesDGX agent

arXiv:2603.12249v2 Announce Type: replace-cross Abstract: Constructing scientific multimodal document reasoning datasets for foundation model training involves an inherent trade-off among scale, faith

SecMate: Multi-Agent Adaptive Cybersecurity Troubleshooting with Tri-Context Personalization

Local AiDGX agent

arXiv:2604.26394v1 Announce Type: cross Abstract: Recent advances in large language models and agentic frameworks have enabled virtual customer assistants (VCAs) for complex support. We present SecMat

Seeking Consensus: Geometric-Semantic On-the-Fly Recalibration for Open-Vocabulary Remote Sensing Semantic Segmentation

SafetyDGX agent

arXiv:2604.26221v1 Announce Type: cross Abstract: Open-vocabulary semantic segmentation (OVSS) in remote sensing images is a promising task that employs textual descriptions for identifying undefined

SG-UniBuc-NLP at SemEval-2026 Task 6: Multi-Head RoBERTa with Chunking for Long-Context Evasion Detection

ResearchDGX agent

arXiv:2604.26375v1 Announce Type: cross Abstract: We describe our system for SemEval-2026 Task 6 (CLARITY: Unmasking Political Question Evasions), which classifies English political interview response

Sociodemographic Biases in Educational Counselling by Large Language Models

SafetyDGX agent

arXiv:2604.25932v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly integrated into educational settings, understanding their potential biases is critical. This study ex

SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment

Model ReleasesDGX agent

arXiv:2604.25937v1 Announce Type: cross Abstract: Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professi

Speech Emotion Recognition Using MFCC Features and LSTM-Based Deep Learning Model

TutorialsDGX agent

arXiv:2604.25938v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) is the use of machines to detect the emotional state of humans based on the speech, which is gaining importance in na

Star-Fusion: A Multi-modal Transformer Architecture for Discrete Celestial Orientation via Spherical Topology

AgentsDGX agent

arXiv:2604.26582v1 Announce Type: cross Abstract: Reliable celestial attitude determination is a critical requirement for autonomous spacecraft navigation, yet traditional 'Lost-in-Space' (LIS) algori

STLGT: A Scalable Trace-Based Linear Graph Transformer for Tail Latency Prediction in Microservices

ApplicationsDGX agent

arXiv:2604.26422v1 Announce Type: cross Abstract: Accurate end-to-end tail-latency forecasting is critical for proactive SLO management in microservice systems. However, modeling long-range dependency

StratMem-Bench: Evaluating Strategic Memory Use in Virtual Character Conversation Beyond Factual Recall

Model ReleasesDGX agent

arXiv:2604.26243v1 Announce Type: cross Abstract: Achieving realistic human-like conversation for virtual characters requires not only a simple memorization and recall of past events, but also the str

Stress Testing Factual Consistency Metrics for Long-Document Summarization

Model ReleasesDGX agent

arXiv:2511.07689v2 Announce Type: replace-cross Abstract: Evaluating the factual consistency of abstractive text summarization remains a significant challenge, particularly for long documents, where c

Structural Generalization on SLOG without Hand-Written Rules

Model ReleasesDGX agent

arXiv:2604.26157v1 Announce Type: cross Abstract: Structural generalization in semantic parsing requires systems to apply learned compositional rules to novel structural combinations. Existing approac

Student Guides Teacher: Weak-to-Strong Inference via Spectral Orthogonal Exploration

SafetyDGX agent

arXiv:2601.06160v2 Announce Type: replace Abstract: Large Language Models (LLMs) often suffer from ''Reasoning Collapse'' on challenging mathematical reasoning tasks, where stochastic sampling produce

SynSur: An end-to-end generative pipeline for synthetic industrial surface defect generation and detection

ResearchDGX agent

arXiv:2604.26633v1 Announce Type: cross Abstract: The bottleneck in learning-based industrial defect detection is often limited not by model capacity, but by the scarcity of labeled defect data: defec

Tatemae: Detecting Alignment Faking via Tool Selection in LLMs

SafetyDGX agent

arXiv:2604.26511v1 Announce Type: cross Abstract: Alignment faking (AF) occurs when an LLM strategically complies with training objectives to avoid value modification, reverting to prior preferences o

TDD Governance for Multi-Agent Code Generation via Prompt Engineering

AgentsDGX agent

arXiv:2604.26615v1 Announce Type: cross Abstract: Large language models (LLMs) accelerate software development but often exhibit instability, non-determinism, and weak adherence to development discipl

Test-Time Safety Alignment

SafetyDGX agent

arXiv:2604.26167v1 Announce Type: cross Abstract: Recent work has shown that a model's input word embeddings can serve as effective control variables for steering its behavior toward outputs that sati

Text Style Transfer with Machine Translation for Graphic Designs

SafetyDGX agent

arXiv:2604.26361v1 Announce Type: cross Abstract: Globalization of graphic designs such as those used in marketing materials and magazines is increasingly important for communication to broad audience

Text-Utilization for Encoder-dominated Speech Recognition Models

ResearchDGX agent

arXiv:2604.26514v1 Announce Type: cross Abstract: This paper investigates efficient methods for utilizing text-only data to improve speech recognition, focusing on encoder-dominated models that facili

The Dual Role of Abstracting over the Irrelevant in Symbolic Explanations: Cognitive Effort vs. Understanding

ResearchDGX agent

arXiv:2602.03467v2 Announce Type: replace Abstract: Explanations are central to human cognition, yet AI systems often produce outputs that are difficult to understand. While symbolic AI offers a trans

The Fools are Certain; the Wise are Doubtful: Exploring LLM Confidence in Code Completion

ResearchDGX agent

arXiv:2508.16131v2 Announce Type: replace-cross Abstract: Code completion entails the task of providing missing tokens given a surrounding context. It can boost developer productivity while providing

TildeOpen LLM: Leveraging Curriculum Learning to Achieve Equitable Language Representation

Model ReleasesDGX agent

arXiv:2603.08182v2 Announce Type: replace-cross Abstract: Large language models often underperform in many European languages due to the dominance of English and a few high-resource languages in train

Time Blindness: Why Video-Language Models Can't See What Humans Can?

Model ReleasesDGX agent

arXiv:2505.24867v2 Announce Type: replace-cross Abstract: Recent advances in vision-language models (VLMs) have made impressive strides in understanding spatio-temporal relationships in videos. Howeve

TimeMM: Time-as-Operator Spectral Filtering for Dynamic Multimodal Recommendation

ApplicationsDGX agent

arXiv:2604.26247v1 Announce Type: cross Abstract: Multimodal recommendation improves user modeling by integrating collaborative signals with heterogeneous item content. In real applications, user inte

TinyR1-32B-Preview: Boosting Accuracy with Branch-Merge Distillation

Model ReleasesDGX agent

arXiv:2503.04872v3 Announce Type: replace-cross Abstract: The challenge of reducing the size of Large Language Models (LLMs) while maintaining their performance has gained significant attention. Howev

TLPO: Token-Level Policy Optimization for Mitigating Language Confusion in Large Language Models

Local AiDGX agent

arXiv:2604.26553v1 Announce Type: cross Abstract: Large language models (LLMs) demonstrate strong multilingual capabilities, yet often fail to consistently generate responses in the intended language,

ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling

ResearchDGX agent

arXiv:2510.14703v2 Announce Type: replace Abstract: Large language models (LLMs) excel at function calling, but inference scaling has been explored mainly for unstructured generation. We propose an in

Training Computer Use Agents to Assess the Usability of Graphical User Interfaces

AgentsDGX agent

arXiv:2604.26020v1 Announce Type: cross Abstract: Usability testing with experts and potential users can assess the effectiveness, efficiency, and user satisfaction of graphical user interfaces (GUIs)

Training-Free Adaptation of New-Generation LLMs using Legacy Clinical Models

Model ReleasesDGX agent

arXiv:2601.03423v3 Announce Type: replace-cross Abstract: Adapting language models to the clinical domain through continued pretraining and instruction tuning requires costly retraining for each new m

Translating Under Pressure: Domain-Aware LLMs for Crisis Communication

SafetyDGX agent

arXiv:2604.26597v1 Announce Type: cross Abstract: Timely and reliable multilingual communication is critical during natural and human-induced disasters, but developing effective solutions for crisis c

Tree-of-Text: A Tree-based Prompting Framework for Table-to-Text Generation in the Sports Domain

ResearchDGX agent

arXiv:2604.26501v1 Announce Type: cross Abstract: Generating sports game reports from structured tables is a complex table-to-text task that demands both precise data interpretation and fluent narrati

Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models

ResearchDGX agent

arXiv:2604.26951v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) offer parallel decoding and bidirectional context, but state-of-the-art dLLMs require billions of parameters f

Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking

SafetyDGX agent

arXiv:2604.26360v1 Announce Type: cross Abstract: Reinforcement learning (RL) systems typically optimize scalar reward functions that assume precise and reliable evaluation of outcomes. However, real-

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising

ResearchDGX agent

arXiv:2604.26694v1 Announce Type: cross Abstract: We propose X-WAM, a Unified 4D World Model that unifies real-time robotic action execution and high-fidelity 4D world synthesis (video + 3D reconstruc

Value-Guided Iterative Refinement and the DIQ-H Benchmark for Evaluating VLM Robustness

Model ReleasesDGX agent

arXiv:2512.03992v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are essential for embodied AI and safety-critical applications, such as robotics and autonomous systems. However

Vertex Features for Neural Global Illumination

ResearchDGX agent

arXiv:2508.07852v2 Announce Type: replace-cross Abstract: Recent research on learnable neural representations has been widely adopted in the field of 3D scene reconstruction and neural rendering appli

Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks

SafetyDGX agent

arXiv:2509.09870v2 Announce Type: replace-cross Abstract: Large language models (LLMs) enable conversational agents (CAs) to express distinctive personalities, raising new questions about how such des

ViCrop-Det: Spatial Attention Entropy Guided Cropping for Training-Free Small-Object Detection

Local AiDGX agent

arXiv:2604.26806v1 Announce Type: cross Abstract: Transformer-based architectures have established a dominant paradigm in global semantic perception; however, they remain fundamentally constrained by

When to Retrieve During Reasoning: Adaptive Retrieval for Large Reasoning Models

Model ReleasesDGX agent

arXiv:2604.26649v1 Announce Type: cross Abstract: Large reasoning models such as DeepSeek-R1 and OpenAI o1 generate extended chains of thought spanning thousands of tokens, yet their integration with

When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling

ResearchDGX agent

arXiv:2604.26644v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on mathematical reasoning tasks but remain unreliable on challenging instances. Existing test-t

Why Attend to Everything? Focus is the Key

ResearchDGX agent

arXiv:2604.03260v2 Announce Type: replace-cross Abstract: Standard attention scales quadratically with sequence length. Efficient attention methods reduce this O(n^2) cost, but when retrofitted into p

28 Apr 2026

A Decoupled Human-in-the-Loop System for Controlled Autonomy in Agentic Workflows

SafetyDGX agent

arXiv:2604.23049v1 Announce Type: new Abstract: AI agents are increasingly deployed to execute tasks and make decisions within agentic workflows, introducing new requirements for safe and controlled a

A Fano-Style Accuracy Upper Bound for LLM Single-Pass Reasoning in Multi-Hop QA

Model ReleasesDGX agent

arXiv:2509.21199v3 Announce Type: replace Abstract: Multi-Hop Question Answering (MHQA) requires integrating dispersed, interdependent evidence through sequential reasoning under noise. This task is c

A General Framework for Generative Self-supervised Learning in Non-invasive Estimation of Physiological Parameters Using Photoplethysmography

Model ReleasesDGX agent

arXiv:2604.22780v1 Announce Type: cross Abstract: Aligning physiological parameter labels with large-scale photoplethysmographic (PPG) data for deep learning is challenging and resource-intensive. Whi

A Lightweight Explainable Guardrail for Prompt Safety

SafetyDGX agent

arXiv:2602.15853v2 Announce Type: replace-cross Abstract: We propose a lightweight explainable guardrail (LEG) method to detect unsafe prompts. LEG uses a multi-task learning architecture to jointly l

A Lower Bound for the Number of Linear Regions of Ternary ReLU Regression Neural Networks

ResearchDGX agent

arXiv:2507.16079v2 Announce Type: replace-cross Abstract: With the advancement of deep learning, reducing computational complexity and memory consumption has become a critical challenge, and ternary n

A Milestone in Formalization: The Sphere Packing Problem in Dimension 8

ResearchDGX agent

arXiv:2604.23468v1 Announce Type: cross Abstract: In 2016, Viazovska famously solved the sphere packing problem in dimension 8, using modular forms to construct a 'magic' function satisfying optimalit

A Parametric Memory Head for Continual Generative Retrieval

Model ReleasesDGX agent

arXiv:2604.23388v1 Announce Type: cross Abstract: Generative information retrieval (GenIR) consolidates retrieval into a single neural model that decodes document identifiers (docids) directly from qu

A Self-Supervised Framework for Space Object Behaviour Characterisation

SafetyDGX agent

arXiv:2504.06176v3 Announce Type: replace-cross Abstract: Foundation Models, which leverage large neural networks pre-trained on unlabelled data before fine-tuning for specific tasks, are increasingly

A Systematic Approach for Large Language Models Debugging

AgentsDGX agent

arXiv:2604.23027v1 Announce Type: new Abstract: Large language models (LLMs) have become central to modern AI workflows, powering applications from open-ended text generation to complex agent-based re

A systematic evaluation of vision-language models for observational astronomical reasoning tasks

Model ReleasesDGX agent

arXiv:2604.24589v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly proposed as general-purpose tools for scientific data interpretation, yet their reliability on real astro

← Previous
1…294295296297298…354
Next →