AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
8 Jun 2026

Queen-Bee Agents: A BeeSpec-Centered Architecture for Governed Enterprise MCP Orchestration

Local AiDGX agent

arXiv:2606.06545v1 Announce Type: cross Abstract: Enterprise agent systems increasingly need to connect large language models to private tools, internal knowledge, and Model Context Protocol (MCP) int

RAVEN: Retrieval-Augmented Vulnerability Exploration Network for Memory Corruption Analysis in User Code and Binary Programs

SafetyDGX agent

arXiv:2604.17948v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities across various cybersecurity tasks, including vulnerability classificat

Re-Centering Humans in LLM Personalization

ResearchDGX agent

arXiv:2606.06614v1 Announce Type: cross Abstract: Despite growing interest, most evaluations of large language models' (LLMs') personalization abilities have relied on synthetic data. It remains uncle


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Re-imagining ISO 26262 in the Age of Autonomous Vehicles: Enhancing Controllability through Transferability and Predictability

SafetyDGX agent

arXiv:2606.07437v1 Announce Type: cross Abstract: The ISO 26262 standard defines functional safety for road vehicles through risk assessments based on Severity, Exposure, and Controllability, grounded

ReclAIm: A Multi-Agent Framework for Monitoring and Correcting Performance Decline in Medical Imaging AI

Model ReleasesDGX agent

arXiv:2510.17004v2 Announce Type: replace-cross Abstract: Purpose: To develop and evaluate a multi-agent framework (ReclAIm) for automated monitoring, detection, and correction of performance decline

REMEDI: A Benchmark for Retention and Unlearning Evaluation in Multi-label Clinical Disease Inference

Model ReleasesDGX agent

arXiv:2606.07141v1 Announce Type: cross Abstract: Language models trained for clinical disease inference are trained on patient data, which may include sensitive and private information, and data owne

RePo: Language Models with Context Re-Positioning

ResearchDGX agent

arXiv:2512.14391v3 Announce Type: replace-cross Abstract: In-context learning is fundamental to modern Large Language Models (LLMs); however, prevailing architectures impose a rigid and fixed contextu

Rethinking Genomic Modeling Through Optical Character Recognition

ResearchDGX agent

arXiv:2602.02014v2 Announce Type: replace-cross Abstract: Recent genomic foundation models largely adopt large language model architectures that treat DNA as a one-dimensional token sequence. However,

RETROSPECT: RETROsynthesis via Sequential Prediction, and Chemically Transformed-ranking

Model ReleasesDGX agent

arXiv:2606.07181v1 Announce Type: cross Abstract: Single-step retrosynthesis needs both accurate first-ranked suggestions and candidate lists that are rich enough for downstream selection. We study th

Robust Driving Control for Autonomous Vehicles: An Intelligent General-sum Constrained Adversarial Reinforcement Learning Approach

SafetyDGX agent

arXiv:2510.09041v3 Announce Type: replace-cross Abstract: Deep reinforcement learning (DRL) has demonstrated remarkable success in developing autonomous driving policies. However, its vulnerability to

SafeGene: Reusable Adapters for Transferable Safety Alignment

SafetyDGX agent

arXiv:2606.06519v1 Announce Type: new Abstract: Open-weight LLMs are increasingly fine-tuned into customized assistants, but downstream fine-tuning can weaken safety alignment and make models more vul

SCALE: Scalable Cross-Attention Learning with Extrapolation for Agentic Workflow Scheduling

AgentsDGX agent

arXiv:2606.06820v1 Announce Type: cross Abstract: Agentic Large Language Model (LLM) systems decompose complex tasks into workflow Directed Acyclic Graphs (DAGs) whose primitives must be scheduled on

ScenicRules: An Autonomous Driving Benchmark with Multi-Objective Specifications and Abstract Scenarios

Model ReleasesDGX agent

arXiv:2602.16073v2 Announce Type: replace-cross Abstract: Developing autonomous driving systems for complex traffic environments requires balancing multiple objectives, such as avoiding collisions, ob

SCOUT: Semantic scene COverage via Uncertainty-guided Traversal

AgentsDGX agent

arXiv:2606.06721v1 Announce Type: cross Abstract: Robots that operate over extended periods should not merely visit space; they should progressively understand it. Yet most 3D scene graph pipelines tr

ShallowBench: Benchmarking Generative Drug Design Models on Shallow-Pocket Targets

Model ReleasesDGX agent

arXiv:2606.06717v1 Announce Type: cross Abstract: While generative AI models have demonstrated remarkable success in structure-based drug design, they predominantly rely on deep binding pockets and st

Should You Use Your Large Language Model to Explore or Exploit?

AgentsDGX agent

arXiv:2502.00225v4 Announce Type: replace-cross Abstract: We evaluate the ability of the current generation of large language models (LLMs) to help a decision-making agent facing an exploration-exploi

SleepExplain: Explainable Non-Rapid Eye Movement and Rapid Eye Movement Sleep Stage Classification from EEG Signal

ResearchDGX agent

arXiv:2606.07351v1 Announce Type: cross Abstract: Classification of sleep stages is one of the most important diagnostic approaches for a variety of sleep-related disorders. Electroencephalography (EE

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating

SafetyDGX agent

arXiv:2606.07074v1 Announce Type: cross Abstract: Deep research agents have demonstrated remarkable capabilities in complex information-seeking tasks, yet this power comes at a steep computational cos

Small Language Model Agents Enable Efficient and High-Quality Knowledge Mining

AgentsDGX agent

arXiv:2510.01427v3 Announce Type: replace Abstract: At the core of Deep Research is knowledge mining, the task of extracting structured information from massive unstructured text in response to user i

Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills

SafetyDGX agent

arXiv:2606.07412v1 Announce Type: cross Abstract: LLM-driven software engineering agents have become a central testbed for real-world language-model capability, yet their training remains limited by t

Sparse Subspace-to-Expert Sharing for Task-Agnostic Continual Learning

Model ReleasesDGX agent

arXiv:2606.07500v1 Announce Type: cross Abstract: Continual learning in Large Language Models (LLMs) is hindered by the plasticity-stability dilemma, where acquiring new capabilities often leads to ca

SpectCount: Spectrotemporal Counting via Synthetic Signals Improves Large Audio Language Models

ApplicationsDGX agent

arXiv:2606.06907v1 Announce Type: cross Abstract: Large audio language models (LALMs) extend large language models with an audio encoder and large-scale audio data. However, the scarcity of high-quali

SS-TPT: Stability and Suitability-Guided Test-Time Prompt Tuning for Adversarially Robust Vision-Language Models

TutorialsDGX agent

arXiv:2606.06943v1 Announce Type: cross Abstract: Vision-language models (VLMs) such as CLIP achieve strong zero-shot recognition but remain highly fragile under adversarial perturbations. Recent test

Stable Reasoning, Unstable Responses: Mitigating LLM Deception via Stability Asymmetry

SafetyDGX agent

arXiv:2603.26846v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) expand in capability and application scope, their trustworthiness becomes critical. A vital risk is intrinsic

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents

Local AiDGX agent

arXiv:2606.07027v1 Announce Type: new Abstract: Reinforcement Learning (RL) has become a promising approach for improving GUI Agents in long-horizon, stochastic digital environments, but trajectory-le

Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning

ApplicationsDGX agent

arXiv:2509.05316v2 Announce Type: replace-cross Abstract: A conventional LLM Unlearning setting consists of two subsets -'forget' and 'retain', with the objectives of removing the undesired knowledge

Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models

SafetyDGX agent

arXiv:2602.02600v3 Announce Type: replace-cross Abstract: Diffusion language models (DLMs) have recently emerged as a competitive alternative to autoregressive (AR) models, offering parallel decoding,

STREAM: Stochastic Riemannian Flow Matching with Anisotropic Decoder for Digital Histopathology Image Generation

Model ReleasesDGX agent

arXiv:2606.07036v1 Announce Type: cross Abstract: Synthetic histopathology image generation addresses critical challenges in computational pathology, including patient privacy and the growing need for

Superintelligent Retrieval Agent: The Next Frontier of Agentic Retrieval

Model ReleasesDGX agent

arXiv:2605.06647v2 Announce Type: replace-cross Abstract: Retrieval-augmented agents are increasingly the interface to large knowledge bases, yet most treat retrieval as a black box: they issue explor

Supervision versus Demonstration-Based In-Context Learning for Multiword Expression Classification

Model ReleasesDGX agent

arXiv:2606.07479v1 Announce Type: cross Abstract: Turkish idiomatic light verb constructions (LVCs) are challenging for multiword expression processing because they often share the same surface form a

SV-Detect: AI-generated Text Detection with Steering Vectors

SafetyDGX agent

arXiv:2606.07313v1 Announce Type: cross Abstract: Detecting machine-generated text is especially difficult under distribution shift, such as transfer across domains, source models, and editing attacks

SW-A^2-Bench: Benchmarking Autonomous Software Agent Generation for Agentic Web

Model ReleasesDGX agent

arXiv:2604.04226v2 Announce Type: replace-cross Abstract: The Agentic Web is emerging as a paradigm in which autonomous software agents interact with online resources and with each other to accomplish

SWE-IF: Aligning Code Evaluation with Human Preference

ResearchDGX agent

arXiv:2510.07315v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have catalyzed vibe coding, where users leverage LLMs to generate and iteratively refine code through natural lan

Synthetic Benchmarks Overstate Forward-Forward Scaling: Real-Data Limits of Layer-Local Training

ResearchDGX agent

arXiv:2606.06539v1 Announce Type: cross Abstract: Forward-Forward (FF) learning [Hinton, 2022] replaces backpropagation with strictly layer-local goodness updates. Recent FF-CNN work has narrowed the

Teaching the Way, Not the Answer: Privileged Tutoring Distillation for Multimodal Policy Optimization

SafetyDGX agent

arXiv:2606.07000v1 Announce Type: new Abstract: Recent post-training methods, particularly Reinforcement Learning with Verifiable Rewards (RLVR), have significantly enhanced the reasoning ability of L

Telling stories, making Hanzi: AI-assisted co-creation with elderly migrants in urban China

ResearchDGX agent

arXiv:2507.01548v3 Announce Type: replace-cross Abstract: This paper explores how older migrants in urban China can record stories that everyday language and design often miss. We ran two co-creation

TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment

Model ReleasesDGX agent

arXiv:2606.07451v1 Announce Type: cross Abstract: Vision-language models such as CLIP are highly useful for diverse tasks due to their shared image-text embedding space. Despite this, the image and te

Textual Supervision Enhances Geospatial Representations in Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.07172v1 Announce Type: cross Abstract: Geospatial understanding is a critical yet underexplored dimension in the development of machine learning systems for tasks such as image geolocation

The Fine-Tuning Trap: Evaluating Negative Transfer and the Role of PEFT in Sub-1B Mathematical Reasoning

Model ReleasesDGX agent

arXiv:2606.06920v1 Announce Type: cross Abstract: Deploying Small Language Models (SLMs) on edge devices requires efficient fine-tuning strategies that adapt models to new tasks without degrading thei

The Geography of Algorithmic Judgment: LLM Intermediaries, Place Identity, and Racial Steering in Housing Search

Local AiDGX agent

arXiv:2606.06694v1 Announce Type: cross Abstract: Large language models (LLMs) are rapidly assuming an intermediary role in housing search through the integration of listing platforms within conversat

The Geometry of Representational Failures in Vision Language Models

Model ReleasesDGX agent

arXiv:2602.07025v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) exhibit puzzling failures in multi-object visual tasks, such as hallucinating non-existent elements or failing t

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

ResearchDGX agent

arXiv:2604.02029v2 Announce Type: replace Abstract: Latent space is rapidly emerging as a native substrate for language-based models. While modern systems are still commonly understood through explici

The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs

Local AiDGX agent

arXiv:2606.07422v1 Announce Type: cross Abstract: Large language models are increasingly used to answer culturally grounded questions across languages, yet it remains unclear whether local cultural kn

The Sim-to-Real Gap of Foundation Model Agents: A Unified MDP Perspective

AgentsDGX agent

arXiv:2606.07017v1 Announce Type: new Abstract: Foundation model agents are increasingly deployed for real-world decision-making, but suffer from the sim-to-real gap. While robotics and classical cont

The Three-Ring Architecture: Governing Agents in the Era of On-Platform Organisations

AgentsDGX agent

arXiv:2606.07119v1 Announce Type: cross Abstract: The current phase of enterprise AI deployment faces a structural failure: organisations are acquiring agentic capability without the infrastructure to

Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models

Model ReleasesDGX agent

arXiv:2606.07157v1 Announce Type: new Abstract: Many efforts to ensure frontier AI models are safe rely on monitoring their chain-of-thought (CoT) reasoning. If models become able to perform sufficien

Think Like a Pilot: Fine-Grained Long-Horizon UAV Navigation

Model ReleasesDGX agent

arXiv:2606.06836v1 Announce Type: cross Abstract: Language-guided UAV agents must execute long-horizon semantic instructions while producing smooth, physically feasible continuous flight commands, yet

ThinkBooster: A Unified Framework for Seamless Test-Time Scaling of LLM Reasoning

Model ReleasesDGX agent

arXiv:2606.06915v1 Announce Type: cross Abstract: Test-time compute (TTC) scaling has emerged as a powerful paradigm for improving large language model (LLM) reasoning by allocating additional compute

TokaMind: A Multi-Modal Transformer Foundation Model for Tokamak Plasma Dynamics

Model ReleasesDGX agent

arXiv:2602.15084v2 Announce Type: replace-cross Abstract: We present TokaMind, to our knowledge the first open-source foundation model for tokamak plasma dynamics, based on a Multi-Modal Transformer (

TOPSIS-RAD: Ranking According to Desires

ResearchDGX agent

arXiv:2606.07253v1 Announce Type: new Abstract: Traditional TOPSIS derives its reference points -- the Positive Ideal Solution (PIS) and Negative Ideal Solution (NIS) -- from the observed alternative

Towards Efficient and Exact Forgetting Services in Pre-Trained-Model-based Continual Learning

ResearchDGX agent

arXiv:2505.12239v2 Announce Type: replace-cross Abstract: In Continual Learning (CL), using a Pre-Trained Model (PTM) as the feature extractor has become a popular practice. Accompanied by analytic cl

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation

TutorialsDGX agent

arXiv:2606.07015v1 Announce Type: cross Abstract: While song generation and singing voice conversion (SVC) have evolved significantly, they have long been developed isolated: the former lacks zero-sho

TRACE: Trajectory Reasoning through Adaptive Cross-Step Evidence Aggregation for LLM Agents

AgentsDGX agent

arXiv:2606.07054v1 Announce Type: cross Abstract: Autonomous LLM agents can pursue hidden malicious objectives through sequences of individually benign actions, making sabotage difficult to detect usi

Trading Engagement for Sustainability: Carbon-Aware Re-ranking for E-commerce Recommendations

Model ReleasesDGX agent

arXiv:2606.04550v1 Announce Type: cross Abstract: E-commerce recommender systems strongly influence which products users consider and purchase, yet sustainability signals such as Product Carbon Footpr

Training for Technology: Adoption and Productive Use of Generative AI in Legal Analysis

ApplicationsDGX agent

arXiv:2603.04982v3 Announce Type: replace-cross Abstract: Can targeted user training unlock the productive potential of generative artificial intelligence in professional settings? We study this quest

TRUE: A Trustworthy Unified Explanation Framework for Large Language Model Reasoning

Local AiDGX agent

arXiv:2602.18905v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated strong capabilities in complex reasoning tasks, yet their decision-making processes remain diff

TSAQA: Time Series Analysis Question And Answering Benchmark

Model ReleasesDGX agent

arXiv:2601.23204v2 Announce Type: replace Abstract: Time series data are integral to critical applications across domains such as finance, healthcare, transportation, and environmental science. While

Twelve quick tips for designing AI-driven HPC workflows

TutorialsDGX agent

arXiv:2606.07491v1 Announce Type: cross Abstract: High-performance computing (HPC) clusters remain the backbone of large-scale scientific computation, traditionally executing deterministic, linear pip

Understanding Generative Recommendation with Semantic IDs from a Model-scaling View

ResearchDGX agent

arXiv:2509.25522v3 Announce Type: replace Abstract: Recent advancements in generative models have allowed the emergence of a promising paradigm for recommender systems (RS), known as Generative Recomm

UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding

Model ReleasesDGX agent

arXiv:2606.07167v1 Announce Type: cross Abstract: Meaningful multilingual evaluation must test models in the target language and educational context. Urdu, spoken by more than 230 million people, lack

← Previous
1…153154155156157…358
Next →