AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
12 May 2026

Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon

Model ReleasesDGX agent

arXiv:2605.09708v1 Announce Type: cross Abstract: We present Metal-Sci, a 10-task benchmark of scientific Apple Silicon Metal compute kernels spanning six optimization regimes (stencils, all-pairs in

Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

Model ReleasesDGX agent

arXiv:2605.10067v1 Announce Type: cross Abstract: Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing ap

mHC-SSM: Manifold-Constrained Hyper-Connections for State Space Language Models with Stream-Specialized Adapters

HardwareDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.08300v1 Announce Type: cross Abstract: Manifold-Constrained Hyper-Connections (mHC) introduce a stability-motivated variant of multi stream residual mixing by constraining residual stream m

Micro-Defects Expose Macro-Fakes: Detecting AI-Generated Images via Local Distributional Shifts

ResearchDGX agent

arXiv:2605.09296v1 Announce Type: cross Abstract: Recent generative models can produce images that appear highly realistic, raising challenges in distinguishing real and AI-generated images. Yet exist

MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph

Model ReleasesDGX agent

arXiv:2605.10120v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) show remarkable potential for scientific reasoning, yet their performance in specialized domains such as micr

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models

SafetyDGX agent

arXiv:2605.08472v1 Announce Type: new Abstract: The effectiveness of Reinforcement Learning (RL) in Large Language Models (LLMs) depends on the nature and diversity of the data used before and during

MIDUS: Memory-Infused Depth Up-Scaling

Model ReleasesDGX agent

arXiv:2512.13751v2 Announce Type: replace-cross Abstract: Expanding pre-trained language models offers a practical way to increase capacity without training larger models from scratch. Depth Up-Scalin

MIND-Skill: Quality-Guaranteed Skill Generation via Multi-Agent Induction and Deduction

AgentsDGX agent

arXiv:2605.08670v1 Announce Type: new Abstract: Large language model (LLM) powered AI agents have emerged as a promising paradigm for autonomous problem-solving, yet they continue to struggle with com

MineEvolve: Self-Evolution with Accumulated Knowledge for Long-Horizon Embodied Minecraft Agents

ResearchDGX agent

arXiv:2603.13131v3 Announce Type: replace Abstract: Long-horizon embodied intelligence requires agents to improve through interaction, not merely to execute plans generated from static goals. A centra

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?

Model ReleasesDGX agent

arXiv:2605.08816v1 Announce Type: new Abstract: In the animal kingdom, mirror self-recognition is a canonical probe of higher-order cognition, emerging only in some species. We ask whether an analogou

Mitigating Many-shot Jailbreak Attacks with One Single Demonstration

SafetyDGX agent

arXiv:2605.08277v1 Announce Type: cross Abstract: Many-shot jailbreaking (MSJ) causes safety-aligned language models to answer harmful queries by preceding them with many harmful question-answer demon

Mitigating Watermark Forgery in Generative Models via Randomized Key Selection

ResearchDGX agent

arXiv:2507.07871v4 Announce Type: replace-cross Abstract: Watermarking enables GenAI providers to verify whether content was generated by their models. A watermark is a hidden signal in the content, w

Mixture of Layers with Hybrid Attention

ResearchDGX agent

arXiv:2605.09516v1 Announce Type: cross Abstract: Standard Mixture-of-Experts (MoE) transformers route tokens to expert subnetworks within each layer, but the layer structure itself remains monolithic

MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection

Model ReleasesDGX agent

arXiv:2605.10833v1 Announce Type: cross Abstract: Industrial anomaly detection is critical for manufacturing quality control, yet existing datasets mainly focus on static images or sparse views, which

Model Merging Scaling Laws in Large Language Models

ResearchDGX agent

arXiv:2509.24244v4 Announce Type: replace Abstract: We study empirical scaling laws for language model merging measured by cross-entropy. Despite its wide practical use, merging lacks a quantitative r

MolRGen: A Training and Evaluation Setting for De Novo Molecular Generation with Reasonning Models

Model ReleasesDGX agent

arXiv:2603.18256v2 Announce Type: replace-cross Abstract: Recent reasoning-based large language models have shown strong performance on tasks with verifiable outcomes, but their use in de novo molecul

MolWorld: Molecule World Models for Actionable Molecular Optimization

Local AiDGX agent

arXiv:2605.08954v1 Announce Type: cross Abstract: Molecular optimization in drug discovery aims to discover molecules with improved target properties, but practical lead optimization often requires mo

MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring

Model ReleasesDGX agent

arXiv:2605.09684v1 Announce Type: cross Abstract: We introduce a red-teaming methodology that exposes harder-to-catch attacks for coding-agent monitors, suggesting that current practices may under-eli

Monocular Biomechanical Tracking of Fingers with Inverse Kinematics to Foundation Models

SafetyDGX agent

arXiv:2605.09258v1 Announce Type: cross Abstract: Accurate hand and finger tracking from video has significant clinical applications for monitoring activities of daily living and measuring range of mo

MoPO: Incorporating Motion Prior for Occluded Human Mesh Recovery

ResearchDGX agent

arXiv:2605.09856v1 Announce Type: cross Abstract: Although recent studies have made remarkable progress in human mesh recovery, they still exhibit limited robustness to occlusions and often produce in

MPerS: Dynamic MLLM MixExperts Perception-Guided Remote Sensing Scene Segmentation

Model ReleasesDGX agent

arXiv:2605.10769v1 Announce Type: cross Abstract: The multimodal fusion of images and scene captions has been extensively explored and applied in various fields. However, when dealing with complex rem

MTA-RL: Robust Urban Driving via Multi-modal Transformer-based 3D Affordances and Reinforcement Learning

SafetyDGX agent

arXiv:2605.10177v1 Announce Type: cross Abstract: Robust urban autonomous driving requires reliable 3D scene understanding and stable decision-making under dense interactions. However, existing end-to

Multi-Armed Bandits With Best-Action Queries

ResearchDGX agent

arXiv:2605.08287v1 Announce Type: cross Abstract: We study multi-armed bandits (MABs) augmented with best-action queries, in which the learner may additionally query an oracle that reveals the best ar

Multi-layer attentive probing improves transfer of audio representations for bioacoustics

SafetyDGX agent

arXiv:2605.10494v1 Announce Type: cross Abstract: Probing heads map the representations learned from audio by a machine learning model to downstream task labels and are a key component in evaluating r

Multi-Level Graph Attention Network Contrastive Learning for Knowledge-Aware Recommendation

ResearchDGX agent

arXiv:2605.08499v1 Announce Type: cross Abstract: In recent years, the use of edge information provided by knowledge graphs together with the advantages of higher-order connectivity in graph neural ne

Multi-Tier Labeling and Physics-Informed Learning for Orbital Anomaly Detection at Scale

Model ReleasesDGX agent

arXiv:2605.09790v1 Announce Type: cross Abstract: Detecting orbital anomalies, such as maneuvers, atmospheric decay, and attitude upsets, across the rapidly growing population of low-Earth-orbit (LEO)

Multimodal Representation Learning Conditioned on Semantic Relations

SafetyDGX agent

arXiv:2508.17497v2 Announce Type: replace-cross Abstract: Multimodal representation learning has been largely driven by contrastive models such as CLIP, which learn a shared embedding space by alignin

MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing

Model ReleasesDGX agent

arXiv:2605.08163v1 Announce Type: cross Abstract: Text-in-image editing has become a key capability for visual content creation, yet existing benchmarks remain overwhelmingly English-centric and often

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation

SafetyDGX agent

arXiv:2511.07833v3 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a standard recipe for post-training LLMs on reasoning tasks, with Group Relat

NaiAD: Initiate Data-Driven Research for LLM Advertising

ResearchDGX agent

arXiv:2605.09918v1 Announce Type: cross Abstract: Reconciling platform revenue with user experience in LLM advertising motivates a data-centric foundation. We introduce NaiAD, the first comprehensive

NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation

Model ReleasesDGX agent

arXiv:2605.10813v1 Announce Type: new Abstract: LLM-powered multi-agent systems can now automate the full research pipeline from ideation to paper writing, but a fundamental question remains: automati

Narrative Landscape: Mapping Narrative Dispositions Across LLMs

ResearchDGX agent

arXiv:2605.08742v1 Announce Type: cross Abstract: This study proposes a quantitative framework for profiling LLM dispositions as stable, model-specific regularities in output under repeated, controlle

Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents

Model ReleasesDGX agent

arXiv:2605.09863v1 Announce Type: cross Abstract: Production LLM coding agents drift over long sessions: they forget user-specified constraints, slip into mistakes the user already flagged, and confab

Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers

SafetyDGX agent

arXiv:2605.09176v1 Announce Type: cross Abstract: Training large language models requires optimization algorithms that are not only statistically effective, but also computationally and memory efficie

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks

Model ReleasesDGX agent

arXiv:2605.10639v1 Announce Type: new Abstract: The rapid adoption of LLMs in both research and industry highlights the challenges of deploying them safely and reveals a gap in the systematic evaluati

NCO: A Versatile Plug-in for Handling Negative Constraints in Decoding

ResearchDGX agent

arXiv:2605.10065v1 Announce Type: cross Abstract: Controlling Large Language Models (LLMs) to prevent the generation of undesirable content, such as profanity and personally identifiable information (

Neural Cluster First, Route Second: One-Shot Capacitated Vehicle Routing via Differentiable Optimal Transport

Model ReleasesDGX agent

arXiv:2605.09301v1 Announce Type: cross Abstract: The Capacitated Vehicle Routing Problem (CVRP) underpins modern last-mile logistics. Current Neural Combinatorial Optimization (NCO) methods construct

Neural Information Causality

Model ReleasesDGX agent

arXiv:2605.09316v1 Announce Type: cross Abstract: Query-separated computation forces a representation to play an operational role: data are encoded before a query is known, and a later decoder can ans

NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims

Model ReleasesDGX agent

arXiv:2605.08192v1 Announce Type: cross Abstract: Frontier AI safety claims - published assertions that a highly capable general-purpose model is below a threshold of concern, adequately mitigated, or

NeuroGAN-3D: Enhancing Intrinsic Functional Brain Networks via High-Fidelity 3D Generative Super-Resolution

Local AiDGX agent

arXiv:2605.08373v1 Announce Type: cross Abstract: Recent advances in neuroimaging have deepened our understanding of the brain's complex functional and structural organization. Among these, functional

Neuroscience-Inspired Analyses of Visual Interestingness in Multimodal Transformers

ResearchDGX agent

arXiv:2605.08188v1 Announce Type: cross Abstract: Human attention is the gateway to conscious perception, memory and decision-making. However, its role in modern transformer models remains largely une

New AI-Driven Tools for Enhancing Campus Well-being: A Prevention and Intervention Approach

Local AiDGX agent

arXiv:2605.10804v1 Announce Type: new Abstract: Campus well-being underpins academic success, yet many universities lack effective methods for monitoring satisfaction and detecting mental health risks

NEXUS: Continual Learning of Symbolic Constraints for Safe and Robust Embodied Planning

SafetyDGX agent

arXiv:2605.09387v1 Announce Type: new Abstract: While Large Language Models (LLMs) have catalyzed progress in embodied intelligence, a fundamental gap between their inherent probabilistic uncertainty

No Mean Feat: Simple, Strong Baselines for Context Compression

ResearchDGX agent

arXiv:2510.20797v2 Announce Type: replace-cross Abstract: Context compression reduces Transformer inference costs by replacing lengthy inputs with shorter pre-computed representations. It carries sign

NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training

ResearchDGX agent

arXiv:2605.08144v1 Announce Type: cross Abstract: Diffusion models have achieved remarkable success across a wide range of generative tasks, yet their training paradigm largely treats injected noise a

NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning

ResearchDGX agent

arXiv:2605.08221v1 Announce Type: cross Abstract: This paper presents NoisyCoconut, a novel inference-time method that enhances large language model (LLM) reliability by manipulating internal represen

Normalization Equivariance for Arbitrary Backbones, with Application to Image Denoising

Model ReleasesDGX agent

arXiv:2605.08193v1 Announce Type: cross Abstract: Normalization Equivariance (NE), equivariance to global contrast and brightness transforms, improves robustness to distribution shift in image-to-imag

Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking

SafetyDGX agent

arXiv:2605.08778v1 Announce Type: new Abstract: Deploying LLMs in multi-turn dialogues facilitates jailbreak attacks that distribute harmful intent across seemingly benign turns. Recent training-based

Not-So-Strange Love: Language Models and Generative Linguistic Theories are More Compatible than They Appear

ResearchDGX agent

arXiv:2605.10061v1 Announce Type: cross Abstract: Futrell and Mahowald (2025) frame the success of neural language models (LMs) as supporting gradient, usage-based linguistic theories. I argue that LM

Novel GPU Boruta algorithms for feature selection from high-dimensional data

HardwareDGX agent

arXiv:2605.09950v1 Announce Type: cross Abstract: Most feature selection algorithms, especially wrapper methods, run inefficiently on CPU based platforms because of their high computational complexity

Octopus Protocol: One-Shot Hardware Discovery and Control for AI Agents via Infrastructure-as-Prompts

HardwareDGX agent

arXiv:2605.09055v1 Announce Type: cross Abstract: Recent agentic-robotics systems, from Code-asPolicies to modern vision-language-action (VLA) foundation models, presuppose that drivers, SDKs, or ROS-

Omni-scale Learning-based Sequential Decision Framework for Order Fulfillment of Tote-handling Robotic Systems

AgentsDGX agent

arXiv:2605.08758v1 Announce Type: cross Abstract: Driven by the rapid expansion of e-commerce and small-batch production, the size of the intralogistics load unit of finished goods, semi-finished good

On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspective

Local AiDGX agent

arXiv:2605.08368v1 Announce Type: new Abstract: Debates about large language model post-training often treat supervised fine-tuning (SFT) as imitation and reinforcement learning (RL) as discovery. But

On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency

ApplicationsDGX agent

arXiv:2601.21619v2 Announce Type: replace-cross Abstract: Parallel thinking improves LLM reasoning through multi-path sampling and aggregation. In standard evaluations, due to a lack of sample-specifi

On Variance Reduction in Learning Mean Flows

SafetyDGX agent

arXiv:2605.09235v1 Announce Type: cross Abstract: One-step generative modeling has emerged as a leading approach to amortize the inference cost of diffusion and flow-matching models. Among distillatio

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.09727v1 Announce Type: cross Abstract: A central challenge in reinforcement learning (RL) is to learn models that generalize beyond the tasks on which they are trained, a goal traditionally

One-Step Graph-Structured Neural Flows for Irregular Multivariate Time Series Classification

TutorialsDGX agent

arXiv:2605.10179v1 Announce Type: cross Abstract: Neural Flows efficiently model irregular multivariate time series by directly learning ODE solution trajectories with neural networks, bypassing step-

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory

ResearchDGX agent

arXiv:2505.23617v3 Announce Type: replace-cross Abstract: Effective video tokenization is critical for scaling transformer models for long videos. Current approaches tokenize videos using space-time p

Open Ontologies: Tool-Augmented Ontology Engineering with Stable Matching Alignment

SafetyDGX agent

arXiv:2605.09184v1 Announce Type: new Abstract: We present Open Ontologies, an open-source ontology engineering system implemented in Rust that integrates LLM-driven construction with formal OWL reaso

OpenClaw-RL: Train Any Agent Simply by Talking

SafetyDGX agent

arXiv:2603.10165v2 Announce Type: replace-cross Abstract: Every agent interaction generates a next-state signal, namely the user reply, tool output, terminal or GUI state change that follows each acti

← Previous
1…270271272273274…358
Next →