AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
23 Apr 2026

CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge

Model ReleasesDGX agent

arXiv:2604.20389v1 Announce Type: cross Abstract: The rapid evolution and use of Large Language Models (LLMs) in professional workflows require an evaluation of their domain-specific knowledge against

DAIRE: A lightweight AI model for real-time detection of Controller Area Network attacks in the Internet of Vehicles

SafetyDGX agent

arXiv:2604.20771v1 Announce Type: cross Abstract: The Internet of Vehicles (IoV) is advancing modern transportation by improving safety, efficiency, and intelligence. However, the reliance on the Cont

Deconstructing Superintelligence: Identity, Self-Modification and Differance

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.19845v1 Announce Type: new Abstract: Self-modification is often taken as constitutive of artificial superintelligence (SI), yet modification is a relative action requiring a supplement outs

Degrees, Levels, and Profiles of Contextuality

ResearchDGX agent

arXiv:2603.26692v3 Announce Type: replace-cross Abstract: We introduce a new notion, that of a contextuality profile of a system of random variables. Rather than characterizing a system's contextualit

Depression Risk Assessment in Social Media via Large Language Models

ResearchDGX agent

arXiv:2604.19887v1 Announce Type: cross Abstract: Depression is one of the most prevalent and debilitating mental health conditions worldwide, frequently underdiagnosed and undertreated. The prolifera

Device-Native Autonomous Agents for Privacy-Preserving Negotiations

Local AiDGX agent

arXiv:2601.00911v3 Announce Type: replace-cross Abstract: Automated negotiations in insurance and business-to-business (B2B) commerce encounter substantial challenges. Current systems force a trade-of

Diagnosing CFG Interpretation in LLMs

SafetyDGX agent

arXiv:2604.20811v1 Announce Type: new Abstract: As LLMs are increasingly integrated into agentic systems, they must adhere to dynamically defined, machine-interpretable interfaces. We evaluate LLMs as

DialToM: A Theory of Mind Benchmark for Forecasting State-Driven Dialogue Trajectories

Model ReleasesDGX agent

arXiv:2604.20443v1 Announce Type: cross Abstract: Large Language Models (LLMs) have been shown to possess Theory of Mind (ToM) abilities. However, it remains unclear whether this stems from robust rea

DISCA: A Digital In-memory Stochastic Computing Architecture Using A Compressed Bent-Pyramid Format

ApplicationsDGX agent

arXiv:2511.17265v2 Announce Type: replace-cross Abstract: Nowadays, we are witnessing an Artificial Intelligence revolution that dominates the technology landscape in various application domains, such

DistortBench: Benchmarking Vision Language Models on Image Distortion Identification

Model ReleasesDGX agent

arXiv:2604.19966v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used in settings where sensitivity to low-level image degradations matters, including content moderatio

Do Hallucination Neurons Generalize? Evidence from Cross-Domain Transfer in LLMs

ApplicationsDGX agent

arXiv:2604.19765v1 Announce Type: cross Abstract: Recent work identifies a sparse set of 'hallucination neurons' (H-neurons), less than 0.1% of feed-forward network neurons, that reliably predict when

Do Small Language Models Know When They're Wrong? Confidence-Based Cascade Scoring for Educational Assessment

Model ReleasesDGX agent

arXiv:2604.19781v1 Announce Type: cross Abstract: Automated scoring of student work at scale requires balancing accuracy against cost and latency. In 'cascade' systems, small language models (LMs) han

Do We Need Bigger Models for Science? Task-Aware Retrieval with Small Language Models

ResearchDGX agent

arXiv:2604.01965v2 Announce Type: replace-cross Abstract: Scientific knowledge discovery increasingly relies on large language models, yet many existing scholarly assistants depend on proprietary syst

DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data

AgentsDGX agent

arXiv:2604.19859v1 Announce Type: cross Abstract: Edge-scale deep research agents based on small language models are attractive for real-world deployment due to their advantages in cost, latency, and

Dual Causal Inference: Integrating Backdoor Adjustment and Instrumental Variable Learning for Medical VQA

Model ReleasesDGX agent

arXiv:2604.20306v1 Announce Type: cross Abstract: Medical Visual Question Answering (MedVQA) aims to generate clinically reliable answers conditioned on complex medical images and questions. However,

Early-Stage Product Line Validation Using LLMs: A Study on Semi-Formal Blueprint Analysis

Model ReleasesDGX agent

arXiv:2604.20523v1 Announce Type: cross Abstract: We study whether Large Language Models (LLMs) can perform feature model analysis operations (AOs) directly on semi-formal textual blueprints, i.e., co

Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models

Model ReleasesDGX agent

arXiv:2511.06209v4 Announce Type: replace Abstract: LLMs can solve complex tasks by generating long, multi-step reasoning chains. Test-time scaling (TTS) can further improve LLM performance by samplin

EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training

SafetyDGX agent

arXiv:2604.20012v1 Announce Type: cross Abstract: Vision-Language-Action Models (VLAs) inherit their visual and linguistic capabilities from Vision-Language Models (VLMs), yet most VLAs are built from

Emergence Transformer: Dynamical Temporal Attention Matters

ResearchDGX agent

arXiv:2604.19816v1 Announce Type: new Abstract: The Transformer, a breakthrough architecture in artificial intelligence, owes its success to the attention mechanism, which utilizes long-range interact

Enhancing ASR Performance in the Medical Domain for Dravidian Languages

TutorialsDGX agent

arXiv:2604.19797v1 Announce Type: cross Abstract: Automatic Speech Recognition (ASR) for low-resource Dravidian languages like Telugu and Kannada faces significant challenges in specialized medical do

Enhancing Research Idea Generation through Combinatorial Innovation and Multi-Agent Iterative Search Strategies

AgentsDGX agent

arXiv:2604.20548v1 Announce Type: cross Abstract: Scientific progress depends on the continual generation of innovative re-search ideas. However, the rapid growth of scientific literature has greatly

Enhancing Speaker Verification with Whispered Speech via Post-Processing

ResearchDGX agent

arXiv:2604.20229v1 Announce Type: cross Abstract: Speaker verification is a task of confirming an individual's identity through the analysis of their voice. Whispered speech differs from phonated spee

Environmental Understanding Vision-Language Model for Embodied Agent

SafetyDGX agent

arXiv:2604.19839v1 Announce Type: cross Abstract: Vision-language models (VLMs) have shown strong perception and reasoning abilities for instruction-following embodied agents. However, despite these a

Epistemic Constitutionalism Or: how to avoid coherence bias

SafetyDGX agent

arXiv:2601.14295v3 Announce Type: replace Abstract: Large language models increasingly function as artificial reasoners: they evaluate arguments, assign credibility, and express confidence. Yet their

Epistemology gives a Future to Complementarity in Human-AI Interactions

SafetyDGX agent

arXiv:2601.09871v2 Announce Type: replace Abstract: Human-AI complementarity is the claim that a human supported by an AI system can outperform either alone in a decision-making process. Since its int

Evian: Towards Explainable Visual Instruction-tuning Data Auditing

Model ReleasesDGX agent

arXiv:2604.20544v1 Announce Type: cross Abstract: The efficacy of Large Vision-Language Models (LVLMs) is critically dependent on the quality of their training data, requiring a precise balance betwee

EvoAgent: An Evolvable Agent Framework with Skill Learning and Multi-Agent Delegation

AgentsDGX agent

arXiv:2604.20133v1 Announce Type: new Abstract: This paper proposes EvoAgent - an evolvable large language model (LLM) agent framework that integrates structured skill learning with a hierarchical sub

EvoForest: A Novel Machine-Learning Paradigm via Open-Ended Evolution of Computational Graphs

Model ReleasesDGX agent

arXiv:2604.19761v1 Announce Type: new Abstract: Modern machine learning is still largely organized around a single recipe: choose a parameterized model family and optimize its weights. Although highly

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2604.19835v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) has become the dominant architecture for scaling large language models: frontier models routinely decouple total parameters f

Explainability in Generative Medical Diffusion Models: A Faithfulness-Based Analysis on MRI Synthesis

ApplicationsDGX agent

arXiv:2602.09781v2 Announce Type: replace-cross Abstract: This study investigates the explainability of generative diffusion models in the context of medical imaging, focusing on Magnetic resonance im

Explainable AML Triage with LLMs: Evidence Retrieval and Counterfactual Checks

SafetyDGX agent

arXiv:2604.19755v1 Announce Type: new Abstract: Anti-money laundering (AML) transaction monitoring generates large volumes of alerts that must be rapidly triaged by investigators under strict audit an

Explainable Speech Emotion Recognition: Weighted Attribute Fairness to Model Demographic Contributions to Social Bias

SafetyDGX agent

arXiv:2604.19763v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) systems have growing applications in sensitive domains such as mental health and education, where biased predictions

Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization

Model ReleasesDGX agent

arXiv:2604.20726v1 Announce Type: cross Abstract: This work explores the role of prompt design and judge selection in LLM-as-a-Judge evaluations of free text legal question answering. We examine wheth

Exploring Data Augmentation and Resampling Strategies for Transformer-Based Models to Address Class Imbalance in AI Scoring of Scientific Explanations in NGSS Classroom

Model ReleasesDGX agent

arXiv:2604.19754v1 Announce Type: new Abstract: Automated scoring of students' scientific explanations offers the potential for immediate, accurate feedback, yet class imbalance in rubric categories p

Fairness Testing of Large Language Models in Role-Playing

Model ReleasesDGX agent

arXiv:2411.00585v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have become foundational in modern language-driven software applications, profoundly influencing daily life. A cr

FeDa4Fair: Client-Level Federated Datasets for Fairness Evaluation

Model ReleasesDGX agent

arXiv:2506.21095v4 Announce Type: replace-cross Abstract: Federated Learning (FL) enables collaborative training while preserving privacy, yet it introduces a critical challenge: the 'illusion of fair

FedSIR: Spectral Client Identification and Relabeling for Federated Learning with Noisy Labels

ResearchDGX agent

arXiv:2604.20825v1 Announce Type: cross Abstract: Federated learning (FL) enables collaborative model training without sharing raw data; however, the presence of noisy labels across distributed client

FELA: A Multi-Agent Evolutionary System for Feature Engineering of Industrial Event Log Data

AgentsDGX agent

arXiv:2510.25223v3 Announce Type: replace Abstract: Event log data, recording fine-grained user actions and system events, represent one of the most valuable assets for modern digital services. Howeve

FLOSS: Federated Learning with Opt-Out and Straggler Support

SafetyDGX agent

arXiv:2507.23115v2 Announce Type: replace-cross Abstract: Previous work on data privacy in federated learning systems focuses on privacy-preserving operations for data from users who have agreed to sh

Forage V2: Knowledge Evolution and Transfer in Autonomous Agent Organizations

AgentsDGX agent

arXiv:2604.19837v1 Announce Type: new Abstract: Autonomous agents operating in open-world tasks -- where the completion boundary is not given in advance -- face denominator blindness: they systematica

Formal Verification of Minimax Algorithms

ResearchDGX agent

arXiv:2509.20138v2 Announce Type: replace Abstract: Minimax-based search algorithms with alpha-beta pruning and transposition tables are a central component of classical game-playing engines and remai

Formalising the Logit Shift Induced by LoRA: A Technical Note

ResearchDGX agent

arXiv:2604.20313v1 Announce Type: cross Abstract: This technical note provides a first-order formalisation of the logit shift and fact-margin change induced by Low-Rank Adaptation (LoRA). Using a firs

Foundation Models in Biomedical Imaging: Turning Hype into Reality

Model ReleasesDGX agent

arXiv:2512.15808v2 Announce Type: replace-cross Abstract: Foundation models (FMs) are driving a prominent shift in biomedical imaging from task-specific models to unified backbone models for diverse t

Frictionless Love: Associations Between AI Companion Roles and Behavioral Addiction

SafetyDGX agent

arXiv:2604.20011v1 Announce Type: cross Abstract: AI companion chatbots increasingly shape how people seek social and emotional connection, sometimes substituting for relationships with romantic partn

From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents

AgentsDGX agent

arXiv:2604.19775v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents capable of reasoning, planning, and acting within interactive environments.

From Admission to Invariants: Measuring Deviation in Delegated Agent Systems

Local AiDGX agent

arXiv:2604.17517v2 Announce Type: replace Abstract: Autonomous agent systems are governed by enforcement mechanisms that flag hard constraint violations at runtime. The Agent Control Protocol identifi

From Data to Theory: Autonomous Large Language Model Agents for Materials Science

Model ReleasesDGX agent

arXiv:2604.19789v1 Announce Type: new Abstract: We present an autonomous large language model (LLM) agent for end-to-end, data-driven materials theory development. The model can choose an equation for

From Fuzzy to Formal: Scaling Hospital Quality Improvement with AI

SafetyDGX agent

arXiv:2604.20055v1 Announce Type: new Abstract: Hospital Quality Improvement (QI) plays a critical role in optimizing healthcare delivery by translating high-level hospital goals into actionable solut

From Noise to Signal to Selbstzweck: Reframing Human Label Variation in the Era of Post-training in NLP

SafetyDGX agent

arXiv:2510.12817v3 Announce Type: replace-cross Abstract: Human Label Variation (HLV) refers to legitimate disagreement in annotation that reflects the diversity of human perspectives rather than mere

From Scene to Object: Text-Guided Dual-Gaze Prediction

Model ReleasesDGX agent

arXiv:2604.20191v1 Announce Type: cross Abstract: Interpretable driver attention prediction is crucial for human-like autonomous driving. However, existing datasets provide only scene-level global gaz

From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization

ResearchDGX agent

arXiv:2604.19884v1 Announce Type: cross Abstract: Post-Training Quantization (PTQ) is critical for the efficient deployment of Large Language Models (LLMs). While 4-bit quantization is widely regarded

FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent Memory

SafetyDGX agent

arXiv:2604.20300v1 Announce Type: new Abstract: For LLM agents, memory management critically impacts efficiency, quality, and security. While much research focuses on retention, selective forgetting--

Generalization and Membership Inference Attack a Practical Perspective

ResearchDGX agent

arXiv:2604.19936v1 Announce Type: cross Abstract: With the emergence of new evaluation metrics and attack methodologies for Membership Inference Attacks (MIA), it becomes essential to reevaluate previ

GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning

SafetyDGX agent

arXiv:2604.20659v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has advanced the reasoning capabilities of Large Language Models (LLMs) by leveraging direct out

Handbook of Rough Set Extensions and Uncertainty Models

ResearchDGX agent

arXiv:2604.19794v1 Announce Type: new Abstract: Rough set theory models uncertainty by approximating target concepts through lower and upper sets induced by indiscernibility, or more generally, by gra

Hidden Reliability Risks in Large Language Models: Systematic Identification of Precision-Induced Output Disagreements

SafetyDGX agent

arXiv:2604.19790v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed under diverse numerical precision configurations, including standard floating-point formats (e.g.

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs

Model ReleasesDGX agent

arXiv:2604.20140v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) is an effective framework for aligning large language models with human preferences, but it struggles with complex

Hybrid Policy Distillation for LLMs

SafetyDGX agent

arXiv:2604.20244v1 Announce Type: cross Abstract: Knowledge distillation (KD) is a powerful paradigm for compressing large language models (LLMs), whose effectiveness depends on intertwined choices of

If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems

AgentsDGX agent

arXiv:2604.19844v1 Announce Type: cross Abstract: Recent advances in embodied Vision-Language Agentic Systems (VLAS), powered by large vision-language models (LVLMs), enable AI systems to perceive and

Image Generators are Generalist Vision Learners

TutorialsDGX agent

arXiv:2604.20329v1 Announce Type: cross Abstract: Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent

← Previous
1…312313314315316…354
Next →