AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
20,981 results
11 Aug 2026

SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning

AgentsDGX agent

arXiv:2505.15062v5 Announce Type: replace-cross Abstract: Knowledge extrapolation is the process of inferring novel information by combining and extending existing knowledge that is explicitly availab

Scale-to-Dialogue: Low-Burden Elicitation of Daily Premenstrual Symptom Ratings with Small Language Models

Local AiDGX agent

arXiv:2608.08746v1 Announce Type: new Abstract: Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal forms impose substantial response burden. We formu

ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.07945v1 Announce Type: cross Abstract: Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource alloc

Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior

ResearchDGX agent

arXiv:2606.22790v2 Announce Type: replace-cross Abstract: In this paper, we investigate the tradeoffs between compute allocation and model performance for two speech processing tasks: Automatic Speech

Scaling Inherently Interpretable Language Models

ResearchDGX agent

arXiv:2608.07594v1 Announce Type: cross Abstract: Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods w

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

Model ReleasesDGX agent

arXiv:2608.09873v1 Announce Type: cross Abstract: We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It co

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning

SafetyDGX agent

arXiv:2608.07959v1 Announce Type: new Abstract: Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multi

SDDBMs: Soft Denoising Diffusion Bridge Models

Model ReleasesDGX agent

arXiv:2608.08594v1 Announce Type: new Abstract: Diffusion bridge models leverage Doob's (h)-transform to construct stochastic transports between arbitrary endpoint distributions, and have shown strong

Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

SafetyDGX agent

arXiv:2608.07531v1 Announce Type: cross Abstract: Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing ext

Second Order Drifting Models

ResearchDGX agent

arXiv:2608.07924v1 Announce Type: cross Abstract: Drifting models are a recent class of one-step generative models that evolve the model distribution during training using a predefined sample-based dr

Second-Order Muon Done Right: A Principled Marriage of Spectral Geometry and Curvature

ResearchDGX agent

arXiv:2608.09763v1 Announce Type: new Abstract: Muon's polar update is exact for an unweighted spectral geometry. We introduce GO-MUON, which uses a matched data-dependent geometry and reuses it acros

See Me, Believe Me: Causality, Intersectionality, and Interventions Improving the Appearance of Patients

ResearchDGX agent

arXiv:2410.01227v2 Announce Type: replace-cross Abstract: In the context of medical records, patients often experience testimonial injustice, where the textual account undermines the validity of their

Self-Attention to Operator Learning-based 3D-IC Thermal Simulation

ResearchDGX agent

arXiv:2510.15968v2 Announce Type: replace-cross Abstract: Thermal management in 3D ICs is increasingly challenging due to higher power densities. Traditional PDE-solving-based methods, while accurate,

Self-Evolving Neuro-Symbolic Skills for Tool-Augmented Spatial Reasoning

AgentsDGX agent

arXiv:2608.07955v1 Announce Type: new Abstract: Large vision-language models have achieved strong performance in multimodal reasoning, but they remain unreliable on fine-grained spatial tasks that dem

SemPIC: Learning Semantic Position-Independent KV Caches

AgentsDGX agent

arXiv:2607.28069v2 Announce Type: replace Abstract: Long-context retrieval and agentic workloads repeatedly reuse the same documents under changing instructions, histories, and document orders. Prefix

Shape Mutating Expert Compression:LorExperts and BTExperts

Model ReleasesDGX agent

arXiv:2608.07814v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many ex

Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic

Model ReleasesDGX agent

arXiv:2601.22510v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often achieve strong benchmark accuracy yet remain brittle under small distribution shifts. While recent mechanis

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

Model ReleasesDGX agent

arXiv:2608.09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, per

Sign Language Recognition Using Original and Synthetic Depth Image Based Point Cloud Data Models

ResearchDGX agent

arXiv:2608.09400v1 Announce Type: cross Abstract: Research regarding the sign language recognition mostly relies on RGB images, whileas sign language datasets that provide depth images are limited. Po

Signature-Guided Capacity Occupancy for Dense Expert Merging

TutorialsDGX agent

arXiv:2608.09201v1 Announce Type: new Abstract: Dense expert merging combines domain-specialized language models into one single checkpoint, typically by admitting task-vector support in weight space.

SignLlama: Enhancing Gloss-free Sign Language Translation by Prioritizing Visual Features for LLMs

TutorialsDGX agent

arXiv:2608.09006v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks. However, fine-tuning LLMs for Gloss-Free Sign Language Tra

SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model

Model ReleasesDGX agent

arXiv:2606.14574v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments. While existing benchmarks

SiriusDeliver: Automating Data Warehouse Delivery at Tencent

AgentsDGX agent

arXiv:2608.09185v1 Announce Type: cross Abstract: Enterprise data warehouses (DWs) support business-critical analytics, but warehouse task delivery remains a complicated production process involving c

SkillConsist: Detecting Inconsistencies in Agent Skills via Bidirectional Graph Alignment

Model ReleasesDGX agent

arXiv:2608.07639v1 Announce Type: cross Abstract: Agent Skills provide reusable capabilities to LLM agents. Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wrong Skill s

SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests

Model ReleasesDGX agent

arXiv:2608.08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the app

SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance

Model ReleasesDGX agent

arXiv:2608.09253v1 Announce Type: new Abstract: LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provide reusable pr

SkillsMetric: Mapping the Detection Boundary of Static Analysis for Malicious Agent Skills

AgentsDGX agent

arXiv:2608.08468v1 Announce Type: cross Abstract: Agent Skills---structured packages of instructions and scripts that augment LLM-based agents---are rapidly proliferating, yet their security propertie

SkillSmith: Enhancing Locally Deployed Agents via Automatic Skill Construction and Evolution

Local AiDGX agent

arXiv:2608.08037v1 Announce Type: new Abstract: LLM-based agent frameworks now act as personal assistants for multi-step tasks. Existing agent frameworks such as OpenClaw commonly follow the Cloud Age

SMAC: Score-Matched Actor-Critics for Robust Offline-to-Online Transfer

SafetyDGX agent

arXiv:2602.17632v3 Announce Type: replace-cross Abstract: Modern offline Reinforcement Learning (RL) methods find performant actor-critics, however, fine-tuning these actor-critics online with value-b

Smart Compaction: Predicting Compaction Utility from Lakehouse Table Metadata

Model ReleasesDGX agent

arXiv:2608.08639v1 Announce Type: new Abstract: Open lakehouse table formats accumulate small data files over time, which degrades query performance. Deciding when compaction is worthwhile remains thr

SNR-Edit: Structure-Aware Noise Rectification for Inversion-Free Flow-Based Editing

ResearchDGX agent

arXiv:2601.19180v2 Announce Type: replace-cross Abstract: Inversion-free image editing using flow-based generative models challenges the prevailing inversion-based pipelines. However, existing approac

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments

Model ReleasesDGX agent

arXiv:2608.09128v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and imp

SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents

Model ReleasesDGX agent

arXiv:2608.08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said. Flat

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation

ResearchDGX agent

arXiv:2608.09271v1 Announce Type: cross Abstract: Group-based reinforcement learning objectives such as GRPO can allocate learning signal poorly across prompt difficulty: under binary rewards, group n

Software Engineering for and with GUI Agent

SafetyDGX agent

arXiv:2608.09278v1 Announce Type: cross Abstract: GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity

SPA: A Simple but Tough-to-Beat Baseline for Knowledge Injection

ResearchDGX agent

arXiv:2603.22213v2 Announce Type: replace-cross Abstract: While large language models (LLMs) are pretrained on massive amounts of data, their knowledge coverage remains incomplete in specialized, data

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention

Model ReleasesDGX agent

arXiv:2608.07921v1 Announce Type: cross Abstract: We apply Marchenko-Pastur (MP) random matrix theory to pre-trained attention weights in order to separate each projection matrix into a random-like bu

SpeedTuning: Speeding Up Policy Execution with Lightweight Reinforcement Learning

SafetyDGX agent

arXiv:2608.09138v1 Announce Type: cross Abstract: While learned robotic policies hold promise for advancing generalizable manipulation, their practical deployment is often hindered by suboptimal execu

SpikeWorld: Fast-State Adaptation for Frozen Spiking World Models

Model ReleasesDGX agent

arXiv:2608.07712v1 Announce Type: cross Abstract: A predictive model receives a self-supervised signal whenever the consequence of an action is observed. Using that signal after deployment is difficul

SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation

Model ReleasesDGX agent

arXiv:2601.09974v2 Announce Type: replace Abstract: Personalizing Large Language Models typically relies on static retrieval or one-time adaptation, assuming user preferences remain invariant over tim

SR-OPSD: Self-Referenced On-Policy Self-Distillation

SafetyDGX agent

arXiv:2608.09745v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on trajectories generated by the policy to be optimized, provi

STAIR: Effective Incident Response Using an End-to-End Agentic Planning Framework

AgentsDGX agent

arXiv:2608.09524v1 Announce Type: cross Abstract: Incident response planning is critical for restoring compromised software systems after cyberattacks. Common practice relies on expert-driven playbook

Stealing Reasoning Traces from Proprietary LLM APIs

SafetyDGX agent

arXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim

STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs

AgentsDGX agent

arXiv:2608.08164v1 Announce Type: cross Abstract: Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured i

Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization

ResearchDGX agent

arXiv:2307.10053v5 Announce Type: replace-cross Abstract: In this paper, we focus on providing convergence guarantees for stochastic subgradient methods in minimizing nonsmooth nonconvex functions. We

StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning

SafetyDGX agent

arXiv:2608.08326v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective approach for improving multimodal reasoning. However, most existing me

Structure-Enhanced Features and Quality-Aware Dynamic Anchor Scoring for Robust Lane Detection

Local AiDGX agent

arXiv:2608.09610v1 Announce Type: cross Abstract: Lane detection requires recovering thin, elongated, and frequently occluded lane structures under challenging driving conditions. While anchor-based d

Structure-Preserving Uncertainty Propagation in First-Order Proof Search

ResearchDGX agent

arXiv:2608.09190v1 Announce Type: new Abstract: GK is a query-directed first-order prover that extends ordinary resolution-based proof search with explicit positive and negative claims, numerical conf

SUM-AgriVLN: Spatial Understanding Memory for Agricultural Vision-and-Language Navigation

Model ReleasesDGX agent

arXiv:2510.14357v2 Announce Type: replace-cross Abstract: Agricultural robots are emerging as powerful assistants across a wide range of agricultural tasks, nevertheless, they are still heavily relyin

SuperCoder: Assembly Program Superoptimization with Large Language Models

Model ReleasesDGX agent

arXiv:2505.11480v4 Announce Type: replace-cross Abstract: Superoptimization is the task of transforming a program into a faster one, and ideally the very fastest possible one, while preserving its inp

SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents

Model ReleasesDGX agent

arXiv:2608.08253v1 Announce Type: new Abstract: AI agents are becoming shared infrastructure, yet durable memory is commonly assembled from separate retrieval, governance, and operational components.

SuperNeuroMAT: An Efficient Matrix-based Simulator for Spiking Neural Networks

ResearchDGX agent

arXiv:2608.08479v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) offer a promising pathway to energy-efficient AI and brain-inspired computing. However, their widespread adoption is hi

SurgLAT: Surgical Latent Attention Tracking for Depth-Aware Robotic Laparoscope Control

AgentsDGX agent

arXiv:2608.07876v1 Announce Type: new Abstract: Autonomous laparoscopic camera control requires continuous understanding of the surgeon's operative intent in dynamic surgical scenes, where the target

Symbolic Attack Chain Generation from Atomic Red Team Techniques: An Empirical Study of Predicate Representation Granularity

SafetyDGX agent

arXiv:2608.00143v2 Announce Type: replace-cross Abstract: Automated attack chain generation is critical for modern cybersecurity, yet manual construction fails to scale as adversary behaviors expand.

SymboUQ: Symbolic Uncertainty Quantification for Spatial Reasoning in LLMs

ResearchDGX agent

arXiv:2608.00417v2 Announce Type: replace Abstract: Although large language models (LLMs) can produce fluent spatial reasoning traces, their intermediate relations may fail to support the final conclu

SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification

Local AiDGX agent

arXiv:2608.08786v1 Announce Type: new Abstract: Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be unfaithful even when final answers ar

Tabular Numeric Stretch Transformation

Model ReleasesDGX agent

arXiv:2608.09162v1 Announce Type: cross Abstract: Tabular data presents unique challenges for deep learning due to its heterogeneous nature, where numeric features exhibit diverse distributions, scale

Targeted Counterfactual Fingerprinting for Black-Box LLM Ownership Verification

Local AiDGX agent

arXiv:2608.08195v1 Announce Type: cross Abstract: Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignment. Because

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

Model ReleasesDGX agent

arXiv:2608.09538v1 Announce Type: cross Abstract: We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation.

TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?

Model ReleasesDGX agent

arXiv:2608.07899v1 Announce Type: new Abstract: Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that failure orig

← Previous
1…1112131415…350
Next →