AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
20,981 results
10 Apr 2026

Restoring Heterogeneity in LLM-based Social Simulation: An Audience Segmentation Approach

Model ReleasesDGX agent

arXiv:2604.06663v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used to simulate social attitudes and behaviors, offering scalable 'silicon samples' that can approximat

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability

SafetyDGX agent

arXiv:2604.06628v1 Announce Type: new Abstract: A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes. We revisit t

REVEAL: Reasoning-Enhanced Forensic Evidence Analysis for Explainable AI-Generated Image Detection

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2511.23158v2 Announce Type: replace-cross Abstract: The rapid progress of visual generative models has made AI-generated images increasingly difficult to distinguish from authentic ones, posing

Riemann-Bench: A Benchmark for Moonshot Mathematics

Model ReleasesDGX agent

arXiv:2604.06802v1 Announce Type: new Abstract: Recent AI systems have achieved gold-medal-level performance on the International Mathematical Olympiad, demonstrating remarkable proficiency at competi

Robustness Risk of Conversational Retrieval: Identifying and Mitigating Noise Sensitivity in Qwen3-Embedding Model

Model ReleasesDGX agent

arXiv:2604.06176v1 Announce Type: cross Abstract: We present an empirical study of embedding-based retrieval under realistic conversational settings, where queries are short, dialogue-like, and weakly

RoSHI: A Versatile Robot-oriented Suit for Human Data In-the-Wild

SafetyDGX agent

arXiv:2604.07331v1 Announce Type: cross Abstract: Scaling up robot learning will likely require human data containing rich and long-horizon interactions in the wild. Existing approaches for collecting

RPM-Net Reciprocal Point MLP Network for Unknown Network Security Threat Detection

TutorialsDGX agent

arXiv:2604.06638v1 Announce Type: cross Abstract: Effective detection of unknown network security threats in multi-class imbalanced environments is critical for maintaining cyberspace security. Curren

$S^3$: Stratified Scaling Search for Test-Time in Diffusion Language Models

ResearchDGX agent

arXiv:2604.06260v1 Announce Type: cross Abstract: Test-time scaling investigates whether a fixed diffusion language model (DLM) can generate better outputs when given more inference compute, without a

SALLIE: Safeguarding Against Latent Language & Image Exploits

Model ReleasesDGX agent

arXiv:2604.06247v1 Announce Type: cross Abstract: Large Language Models (LLMs) and Vision-Language Models (VLMs) remain highly vulnerable to textual and visual jailbreaks, as well as prompt injections

Say Something Else: Rethinking Contextual Privacy as Information Sufficiency

ResearchDGX agent

arXiv:2604.06409v1 Announce Type: cross Abstract: LLM agents increasingly draft messages on behalf of users, yet users routinely overshare sensitive information and disagree on what counts as private.

Scientific Knowledge-driven Decoding Constraints Improving the Reliability of LLMs

ResearchDGX agent

arXiv:2604.06603v1 Announce Type: cross Abstract: Large language models (LLMs) have shown strong knowledge reserves and task-solving capabilities, but still face the challenge of severe hallucination,

SE-Enhanced ViT and BiLSTM-Based Intrusion Detection for Secure IIoT and IoMT Environments

Model ReleasesDGX agent

arXiv:2604.06254v1 Announce Type: cross Abstract: With the rapid growth of interconnected devices in Industrial and Medical Internet of Things (IIoT and MIoT) ecosystems, ensuring timely and accurate

Self-Discovered Intention-aware Transformer for Multi-modal Vehicle Trajectory Prediction

AgentsDGX agent

arXiv:2604.07126v1 Announce Type: cross Abstract: Predicting vehicle trajectories plays an important role in autonomous driving and ITS applications. Although multiple deep learning algorithms are dev

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models

Model ReleasesDGX agent

arXiv:2604.06996v1 Announce Type: cross Abstract: LLM-as-a-judge has become the de facto approach for evaluating LLM outputs. However, judges are known to exhibit self-preference bias (SPB): they tend

Self-Supervised Foundation Model for Calcium-imaging Population Dynamics

Model ReleasesDGX agent

arXiv:2604.04958v2 Announce Type: replace-cross Abstract: Recent work suggests that large-scale, multi-animal modeling can significantly improve neural recording analysis. However, for functional calc

SELFDOUBT: Uncertainty Quantification for Reasoning LLMs via the Hedge-to-Verify Ratio

ApplicationsDGX agent

arXiv:2604.06389v1 Announce Type: new Abstract: Uncertainty estimation for reasoning language models remains difficult to deploy in practice: sampling-based methods are computationally expensive, whil

SensorPersona: An LLM-Empowered System for Continual Persona Extraction from Longitudinal Mobile Sensor Streams

AgentsDGX agent

arXiv:2604.06204v1 Announce Type: cross Abstract: Personalization is essential for Large Language Model (LLM)-based agents to adapt to users' preferences and improve response quality and task performa

SentinelSphere: Integrating AI-Powered Real-Time Threat Detection with Cybersecurity Awareness Training

Model ReleasesDGX agent

arXiv:2604.06900v1 Announce Type: cross Abstract: The field of cybersecurity is confronted with two interrelated challenges: a worldwide deficit of qualified practitioners and ongoing human-factor wea

Severity-Aware Weighted Loss for Arabic Medical Text Generation

Model ReleasesDGX agent

arXiv:2604.06346v1 Announce Type: cross Abstract: Large language models have shown strong potential for Arabic medical text generation; however, traditional fine-tuning objectives treat all medical ca

ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference

Local AiDGX agent

arXiv:2508.16703v4 Announce Type: replace-cross Abstract: On-device running Large Language Models (LLMs) is nowadays a critical enabler towards preserving user privacy. We observe that the attention o

SHAPE: Stage-aware Hierarchical Advantage via Potential Estimation for LLM Reasoning

ResearchDGX agent

arXiv:2604.06636v1 Announce Type: cross Abstract: Process supervision has emerged as a promising approach for enhancing LLM reasoning, yet existing methods fail to distinguish meaningful progress from

SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills

Model ReleasesDGX agent

arXiv:2604.06550v1 Announce Type: cross Abstract: OpenClaw's ClawHub marketplace hosts over 13,000 community-contributed agent skills, and between 13% and 26% of them contain security vulnerabilities

SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems

Model ReleasesDGX agent

arXiv:2604.06811v1 Announce Type: cross Abstract: Skill-based agent systems tackle complex tasks by composing reusable skills, improving modularity and scalability while introducing a largely unexamin

SleepNet and DreamNet: Enriching and Reconstructing Representations for Consolidated Visual Classification

ResearchDGX agent

arXiv:2409.01633v4 Announce Type: replace-cross Abstract: An effective integration of rich feature representations with robust classification mechanisms remains a key challenge in visual understanding

Soft-Quantum Algorithms

SafetyDGX agent

arXiv:2604.06523v1 Announce Type: cross Abstract: Quantum operations on pure states can be fully represented by unitary matrices. Variational quantum circuits, also known as quantum neural networks, e

Space Filling Curves is All You Need: Communication-Avoiding Matrix Multiplication Made Simple

Local AiDGX agent

arXiv:2601.16294v2 Announce Type: replace-cross Abstract: General Matrix Multiplication (GEMM) is the cornerstone of HPC workloads and Deep Learning. State-of-the-art vendor libraries tune tensor layo

Sparse-Aware Neural Networks for Nonlinear Functionals: Mitigating the Exponential Dependence on Dimension

ResearchDGX agent

arXiv:2604.06774v1 Announce Type: cross Abstract: Deep neural networks have emerged as powerful tools for learning operators defined over infinite-dimensional function spaces. However, existing theori

SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization

Model ReleasesDGX agent

arXiv:2511.11663v2 Announce Type: replace-cross Abstract: The emergence of accurate open large language models (LLMs) has sparked a push for advanced quantization techniques to enable efficient deploy

Spectral Edge Dynamics Reveal Functional Modes of Learning

Model ReleasesDGX agent

arXiv:2604.06256v1 Announce Type: cross Abstract: Training dynamics during grokking concentrate along a small number of dominant update directions -- the spectral edge -- which reliably distinguishes

SPICE: Submodular Penalized Information-Conflict Selection for Efficient Large Language Model Training

ResearchDGX agent

arXiv:2601.23155v2 Announce Type: replace-cross Abstract: Information-based data selection for instruction tuning is compelling: maximizing the log-determinant of the Fisher information yields a monot

Stabilizing Unsupervised Self-Evolution of MLLMs via Continuous Softened Retracing reSampling

ResearchDGX agent

arXiv:2604.03647v2 Announce Type: replace-cross Abstract: In the unsupervised self-evolution of Multimodal Large Language Models, the quality of feedback signals during post-training is pivotal for st

Steering the Verifiability of Multimodal AI Hallucinations

TutorialsDGX agent

arXiv:2604.06714v1 Announce Type: new Abstract: AI applications driven by multimodal large language models (MLLMs) are prone to hallucinations and pose considerable risks to human users. Crucially, su

Strategic Persuasion with Trait-Conditioned Multi-Agent Systems for Iterative Legal Argumentation

Model ReleasesDGX agent

arXiv:2604.07028v1 Announce Type: cross Abstract: Strategic interaction in adversarial domains such as law, diplomacy, and negotiation is mediated by language, yet most game-theoretic models abstract

Stress Estimation in Elderly Oncology Patients Using Visual Wearable Representations and Multi-Instance Learning

ResearchDGX agent

arXiv:2604.06990v1 Announce Type: cross Abstract: Psychological stress is clinically relevant in cardio-oncology, yet it is typically assessed only through patient-reported outcome measures (PROMs) an

STRIDE-ED: A Strategy-Grounded Stepwise Reasoning Framework for Empathetic Dialogue Systems

ResearchDGX agent

arXiv:2604.07100v1 Announce Type: cross Abstract: Empathetic dialogue requires not only recognizing a user's emotional state but also making strategy-aware, context-sensitive decisions throughout resp

SubFLOT: Submodel Extraction for Efficient and Personalized Federated Learning via Optimal Transport

Local AiDGX agent

arXiv:2604.06631v1 Announce Type: cross Abstract: Federated Learning (FL) enables collaborative model training while preserving data privacy, but its practical deployment is hampered by system and sta

SurFITR: A Dataset for Surveillance Image Forgery Detection and Localisation

Local AiDGX agent

arXiv:2604.07101v1 Announce Type: cross Abstract: We present the Surveillance Forgery Image Test Range (SurFITR), a dataset for surveillance-style image forgery detection and localisation, in response

SymptomWise: A Deterministic Reasoning Layer for Reliable and Efficient AI Systems

SafetyDGX agent

arXiv:2604.06375v1 Announce Type: new Abstract: AI-driven symptom analysis systems face persistent challenges in reliability, interpretability, and hallucination. End-to-end generative approaches ofte

Syntax Is Easy, Semantics Is Hard: Evaluating LLMs for LTL Translation

ResearchDGX agent

arXiv:2604.07321v1 Announce Type: cross Abstract: Propositional Linear Temporal Logic (LTL) is a popular formalism for specifying desirable requirements and security and privacy policies for software,

Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity

Model ReleasesDGX agent

arXiv:2509.09794v4 Announce Type: replace Abstract: Computational models have emerged as powerful tools for multi-scale energy modeling research at the building and urban scale, supporting data-driven

TalkLoRA: Communication-Aware Mixture of Low-Rank Adaptation for Large Language Models

Model ReleasesDGX agent

arXiv:2604.06291v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning of Large Language Models (LLMs), and recent Mixture-of-Experts (MoE) extensions fur

TeaLeafVision: An Explainable and Robust Deep Learning Framework for Tea Leaf Disease Classification

ApplicationsDGX agent

arXiv:2604.07182v1 Announce Type: cross Abstract: As the worlds second most consumed beverage after water, tea is not just a cultural staple but a global economic force of profound scale and influence

Team Fusion@ SU@ BC8 SympTEMIST track: transformer-based approach for symptom recognition and linking

ResearchDGX agent

arXiv:2604.06424v1 Announce Type: cross Abstract: This paper presents a transformer-based approach to solving the SympTEMIST named entity recognition (NER) and entity linking (EL) tasks. For NER, we f

TeamLLM: A Human-Like Team-Oriented Collaboration Framework for Multi-Step Contextualized Tasks

Model ReleasesDGX agent

arXiv:2604.06765v1 Announce Type: cross Abstract: Recently, multi-Large Language Model (LLM) frameworks have been proposed to solve contextualized tasks. However, these frameworks do not explicitly em

Temporal Inversion for Learning Interval Change in Chest X-Rays

SafetyDGX agent

arXiv:2604.04563v2 Announce Type: replace-cross Abstract: Recent advances in vision--language pretraining have enabled strong medical foundation models, yet most analyze radiographs in isolation, over

Temporally Phenotyping GLP-1RA Case Reports with Large Language Models: A Textual Time Series Corpus and Risk Modeling

Model ReleasesDGX agent

arXiv:2604.06197v1 Announce Type: cross Abstract: Type 2 diabetes case reports describe complex clinical courses, but their timelines are often expressed in language that is difficult to reuse in long

The AI Skills Shift: Mapping Skill Obsolescence, Emergence, and Transition Pathways in the LLM Era

Model ReleasesDGX agent

arXiv:2604.06906v1 Announce Type: cross Abstract: As Large Language Models reshape the global labor market, policymakers and workers need empirical data on which occupational skills may be most suscep

The Art of Building Verifiers for Computer Use Agents

AgentsDGX agent

arXiv:2604.06240v1 Announce Type: cross Abstract: Verifying the success of computer use agent (CUA) trajectories is a critical challenge: without reliable verification, neither evaluation nor training

The ATOM Report: Measuring the Open Language Model Ecosystem

Model ReleasesDGX agent

arXiv:2604.07190v1 Announce Type: cross Abstract: We present a comprehensive adoption snapshot of the leading open language models and who is building them, focusing on the ~1.5K mainline open models

The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?

SafetyDGX agent

arXiv:2604.06436v2 Announce Type: cross Abstract: We prove that no continuous, utility-preserving wrapper defense-a function D: Xo X that preprocesses inputs before the model sees them-can make al

The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning

Model ReleasesDGX agent

arXiv:2604.06427v1 Announce Type: cross Abstract: The viability of chain-of-thought (CoT) monitoring hinges on models being unable to reason effectively in their latent representations. Yet little is

The Detection-Extraction Gap: Models Know the Answer Before They Can Say It

ResearchDGX agent

arXiv:2604.06613v2 Announce Type: cross Abstract: Modern reasoning models continue generating long after the answer is already determined. Across five model configurations, two families, and three ben

The End of the Foundation Model Era: Open-Weight Models, Sovereign AI, and Inference as Infrastructure

SafetyDGX agent

arXiv:2604.06217v1 Announce Type: cross Abstract: The foundation model era -- roughly 2020 to 2025 -- is over. The forces that defined it have inverted. Open source models have reached frontier perfor

The Geometry of Forgetting

Model ReleasesDGX agent

arXiv:2604.06222v1 Announce Type: cross Abstract: Why do we forget? Why do we remember things that never happened? The conventional answer points to biological hardware. We propose a different one: ge

The Human Condition as Reflected in Contemporary Large Language Models

ResearchDGX agent

arXiv:2604.06206v1 Announce Type: cross Abstract: This study seeks to uncover evidence of a latent structure in evolved human culture as it is refracted through contemporary large language models (LLM

The Impact of Steering Large Language Models with Persona Vectors in Educational Applications

Model ReleasesDGX agent

arXiv:2604.07102v1 Announce Type: cross Abstract: Activation-based steering can personalize large language models at inference time, but its effects in educational settings remain unclear. We study pe

The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment

SafetyDGX agent

arXiv:2604.06377v1 Announce Type: cross Abstract: We investigate whether post-trained capabilities can be transferred across models without retraining, with a focus on transfer across different model

The Planetary Cost of AI Acceleration, Part II: The 10th Planetary Boundary and the 6.5-Year Countdown

AgentsDGX agent

arXiv:2604.04956v2 Announce Type: replace-cross Abstract: The recent, super-exponential scaling of autonomous Large Language Model (LLM) agents signals a broader, fundamental paradigm shift from machi

The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?

Model ReleasesDGX agent

arXiv:2604.06192v1 Announce Type: cross Abstract: Recent work uses entropy-based signals at multiple representation levels to study reasoning in large language models, but the field remains largely em

The Traveling Thief Problem with Time Windows: Benchmarks and Heuristics

Model ReleasesDGX agent

arXiv:2604.06724v1 Announce Type: cross Abstract: While traditional optimization problems were often studied in isolation, many real-world problems today require interdependence among multiple optimiz

← Previous
1…347348349350
Next →