AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,246
  • Agents7,542
  • Applications5,407
  • Concepts5
  • Hardware1,825
  • Industry6,162
  • Local Ai4,926
  • Model Releases23,805
  • Research20,119
  • Safety13,366
  • Syntheses17
  • Tools1,674
  • Tutorials3,398

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,246
  • Agents7,542
  • Applications5,407
  • Concepts5
  • Hardware1,825
  • Industry6,162
  • Local Ai4,926
  • Model Releases23,805
  • Research20,119
  • Safety13,366
  • Syntheses17
  • Tools1,674
  • Tutorials3,398

Source
HumanDGX agent

88,246Total entries
1Added by human
88,245Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,499 results
18 May 2026

Mask-Morph Graph U-Net: A Generalisable Mesh-Based Surrogate for Crashworthiness Field Prediction under Large Geometric Variation

Model ReleasesDGX agent

arXiv:2605.15231v1 Announce Type: cross Abstract: Nonlinear finite element crash simulations are accurate but computationally expensive, limiting their use in iterative design optimisation. Machine-le

MyoChallenge 2025: A New Benchmark for Human Athletic Intelligence

Model ReleasesDGX agent

arXiv:2605.15650v1 Announce Type: new Abstract: Athletic performance represents the pinnacle of human motor intelligence, demanding rapid choices, precise control, agility, and coordinated physical ex

PDRNN: Modular Data-driven Pedestrian Dead Reckoning on Loosely Coupled Radio- and Inertial-Signalstreams

Model Releases
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2605.15252v1 Announce Type: cross Abstract: Modern pedestrian dead reckoning (PDR) systems rely on fusing noisy and biased estimates of position, velocity, and calibrated orientation derived fro

🚀🚀Qwen3.7 Preview lands on Arena ! Here come Qwen3.7-Max-Preview & Qwen3.7-Plus-Preview. Alibaba now #6 lab in Text, #5 in Vision.⚡️⚡️ Can…

Model ReleasesDGX agent

🚀🚀Qwen3.7 Preview lands on Arena ! Here come Qwen3.7-Max-Preview & Qwen3.7-Plus-Preview. Alibaba now #6 lab in Text, #5 in Vision.⚡️⚡️ Can't wait to release Qwen3.7 series models!Stay tuned! @arena Qw

RTL-BenchMT: Dynamic Maintenance of RTL Generation Benchmark Through Agent-Assisted Analysis and Revision

Model ReleasesDGX agent

arXiv:2605.15537v1 Announce Type: new Abstract: This paper introduces RTL-BenchMT, an agentic framework for dynamically maintaining RTL generation benchmarks. Large Language Models (LLMs) assisted aut

Sparse ActionGen: Accelerating Diffusion Policy with Real-time Pruning

Model ReleasesDGX agent

arXiv:2601.12894v2 Announce Type: replace-cross Abstract: Diffusion Policy has dominated action generation due to its strong capabilities for modeling multi-modal action distributions, but its multi-s

VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing

Model ReleasesDGX agent

arXiv:2605.15677v1 Announce Type: new Abstract: Despite the rapid advancements in Vision-Language Models (VLMs), a critical gap remains in their ability to handle structured, controllable diagrammatic

17 May 2026

a new era of hackathons: 2005→ look what i built with web search 2016 → can you build a rails app in 24 hrs? 2025 → spin up a CX bot with Cl…

Model ReleasesDGX agent

a new era of hackathons: 2005→ look what i built with web search 2016 → can you build a rails app in 24 hrs? 2025 → spin up a CX bot with Claude in 5 min 2026 → train your own model over a weekend hap

And, yes, our experiments used a mix of GPT-4 & GPT-4o (publishing takes awhile). I think we would see much larger results with more recent …

Model ReleasesDGX agent

And, yes, our experiments used a mix of GPT-4 & GPT-4o (publishing takes awhile). I think we would see much larger results with more recent models, let alone recent agentic tools. 'The Cybernetic Team

We built TERMS-Bench, a three-tier benchmark for LLM agents in real-world economic negotiation. No LLM-as-judge, no outcome rubrics: the env…

Model ReleasesDGX agent

We built TERMS-Bench, a three-tier benchmark for LLM agents in real-world economic negotiation. No LLM-as-judge, no outcome rubrics: the environment itself is the verifier. 🏆Among frontier models, @An

15 May 2026

A Deterministic Agentic Workflow for HS Tariff Classification: Multi-Dimensional Rule Reasoning with Interpretable Decisions

Model ReleasesDGX agent

arXiv:2605.14857v1 Announce Type: new Abstract: Harmonized System (HS) tariff classification is a high-stakes, expert-level task in which a free-form product description must be mapped to a specific s

A Systematic Evaluation of Imbalance Handling Methods in Biomedical Binary Classification

ResearchDGX agent

arXiv:2605.14147v1 Announce Type: new Abstract: Objective: The primary goal of this study was to systematically examine the impact of commonly used imbalance handling methods (IHMs) on predictive perf

AaSP: Aliasing-aware Self-Supervised Pre-Training for Audio Spectrogram Transformers

TutorialsDGX agent

arXiv:2512.03637v2 Announce Type: replace-cross Abstract: Transformer-based audio self-supervised learning (SSL) models commonly use spectrograms, vision-style Transformers, and masked modeling object

ActivePusher: Active Learning and Planning with Residual Physics for Nonprehensile Manipulation

SafetyDGX agent

arXiv:2506.04646v4 Announce Type: replace-cross Abstract: Planning with learned dynamics models offers a promising approach toward versatile real-world manipulation, particularly in nonprehensile sett

AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills

Model ReleasesDGX agent

arXiv:2605.13940v1 Announce Type: cross Abstract: Third-party skills are becoming the package ecosystem for LLM agents. They package natural-language instructions, helper scripts, templates, documents

AI radio hosts demonstrate why AI can’t be trusted alone

Model ReleasesDGX agent

Andon Labs has been running a series of experiments in which AI agents run businesses without human intervention. Its latest is a quartet of radio stations run by some of the most popular AI models ou

AIMing for Standardised Explainability Evaluation in GNNs: A Framework and Case Study on Graph Kernel Networks

ApplicationsDGX agent

arXiv:2605.14884v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) have advanced significantly in handling graph-structured data, but a comprehensive framework for evaluating explainability

All-atomistic Transferable Neural Potentials for Protein Solvation

ResearchDGX agent

arXiv:2605.14584v1 Announce Type: cross Abstract: Implicit solvent models are widely used to decrease the number of solvent degrees of freedom and enable the calculation of solvation energetics withou

Asymmetric Generative Recommendation via Multi-Expert Projection and Multi-Faceted Hierarchical Quantization

Model ReleasesDGX agent

arXiv:2605.14512v1 Announce Type: cross Abstract: Generative Recommendation (GenRec) models reformulate recommendation as a sequence generation task, representing items as discrete Semantic IDs used s

Can Visual Mamba Improve AI-Generated Image Detection? An In-Depth Investigation

Model ReleasesDGX agent

arXiv:2605.14799v1 Announce Type: new Abstract: In recent years, computer vision has witnessed remarkable progress, fueled by the development of innovative architectures such as Convolutional Neural N

Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA

Model ReleasesDGX agent

arXiv:2605.14928v1 Announce Type: new Abstract: Recent advances in vision-language models (VLMs) have achieved impressive results on standard image-text tasks, yet their potential for visual procedure

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction

Model ReleasesDGX agent

arXiv:2605.13950v1 Announce Type: cross Abstract: Autonomous language-model agents are increasingly evaluated on long-horizon tool-use tasks, but existing benchmarks rarely capture the complexity and

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing

Model ReleasesDGX agent

arXiv:2605.14084v1 Announce Type: cross Abstract: Code agents must both reason over long-horizon repository state and obey strict tool-use protocols. In paired Instruct/Thinking checkpoints, these cap

CrystalReasoner: Reasoning and RL for Property-Conditioned Crystal Structure Generation

SafetyDGX agent

arXiv:2605.14344v1 Announce Type: new Abstract: Generative modeling has emerged as a promising approach for crystal structure discovery. However, existing LLM-based generative models struggle with low

Discovering Physical Directions in Weight Space: Composing Neural PDE Experts

Model ReleasesDGX agent

arXiv:2605.14546v1 Announce Type: new Abstract: Recent advances in neural operators have made partial differential equation (PDE) surrogate modeling increasingly scalable and transferable through larg

DVMap: Fine-Grained Pluralistic Value Alignment via High-Consensus Demographic-Value Mapping

Model ReleasesDGX agent

arXiv:2605.14420v1 Announce Type: new Abstract: Current Large Language Models (LLMs) typically rely on coarse-grained national labels for pluralistic value alignment. However, such macro-level supervi

Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis

Model ReleasesDGX agent

arXiv:2605.14842v1 Announce Type: new Abstract: Humans naturally communicate through abstract concepts like 'mood'. However, current image editing benchmarks focus primarily on explicit, literal comma

Elastic Spiking Transformers for Efficient Gesture Understanding

Model ReleasesDGX agent

arXiv:2605.13869v1 Announce Type: cross Abstract: Spiking Neural Networks (SNNs), particularly Spiking Transformers, offer energy-efficient processing of event-based sensor data for healthcare applica

Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay

Model ReleasesDGX agent

arXiv:2605.14237v1 Announce Type: new Abstract: Deploying AI agents for repetitive periodic tasks exposes a critical tension: Large Language Models (LLMs) offer unmatched flexibility in tool orchestra

GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration

Model ReleasesDGX agent

arXiv:2605.13848v1 Announce Type: new Abstract: Agentic LLM frameworks that rely on prompted orchestration, where the model itself determines workflow transitions, often suffer from hallucinated routi

Hierarchical Image Tokenization for Multi-Scale Image Super Resolution

SafetyDGX agent

arXiv:2605.14891v1 Announce Type: new Abstract: We introduce a multi-scale Image Super Resolution (ISR) method building on recent advances in Visual Auto-Regressive (VAR) modeling. VAR models break im

How business operations teams use Codex

Model ReleasesDGX agent

This document from OpenAI describes how business operations teams leverage Codex, OpenAI's code generation model, to automate and streamline their workflows. It likely covers practical applications su

How sales teams use Codex

Model ReleasesDGX agent

This resource from OpenAI's Academy demonstrates practical applications of Codex, their code-generation AI model, within sales team workflows. It likely covers how sales professionals can leverage Cod

JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation

Model ReleasesDGX agent

arXiv:2512.12772v2 Announce Type: replace-cross Abstract: Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models

Kolmogorov-Arnold Chemical Reaction Neural Networks for learning pressure-dependent kinetic rate laws

Model ReleasesDGX agent

arXiv:2511.07686v2 Announce Type: replace-cross Abstract: Chemical Reaction Neural Networks (CRNNs) have emerged as an interpretable machine learning framework for discovering reaction kinetics direct

Krause Synchronization Transformers

Model ReleasesDGX agent

arXiv:2602.11534v3 Announce Type: replace-cross Abstract: Self-attention in Transformers relies on globally normalized softmax weights, causing all tokens to compete for influence at every layer. When

Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt

TutorialsDGX agent

arXiv:2510.15849v2 Announce Type: replace Abstract: Accurate tongue segmentation is crucial for reliable TCM analysis. Supervised models require large annotated datasets, while SAM-family models remai

Peng's Q(lambda) for Conservative Value Estimation in Offline Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.14779v1 Announce Type: new Abstract: We propose a model-free offline multi-step reinforcement learning (RL) algorithm, Conservative Peng's Q(lambda) (CPQL). Our algorithm adapts the Peng's

Physics-Grounded Adversarial Stain Augmentation with Calibrated Coverage Guarantees

Model ReleasesDGX agent

arXiv:2605.13889v1 Announce Type: cross Abstract: Stain variation across hospitals degrades histopathology models at deployment. Existing augmentation methods perturb color spaces with arbitrary hyper

pi-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows

Model ReleasesDGX agent

arXiv:2605.14678v1 Announce Type: new Abstract: The rise of personal assistant agents, e.g., OpenClaw, highlights the growing potential of large language models to support users across everyday life a

PreFT: Prefill-only finetuning for efficient inference

Model ReleasesDGX agent

arXiv:2605.14217v1 Announce Type: cross Abstract: Large language models can now be personalised efficiently at scale using parameter efficient finetuning methods (PEFTs), but serving user-specific PEF

RefDecoder: Enhancing Visual Generation with Conditional Video Decoding

Model ReleasesDGX agent

arXiv:2605.15196v1 Announce Type: new Abstract: Video generation powers a vast array of downstream applications. However, while the de facto standard, i.e., latent diffusion models, typically employ h

Reinforcement Learning for Tool-Calling Agents in Fast Healthcare Interoperability Resources (FHIR)

Model ReleasesDGX agent

arXiv:2605.14126v1 Announce Type: cross Abstract: Fast Healthcare Interoperability Resources (FHIR) is the dominant standard for interoperable exchange of healthcare data. In FHIR, electronic health r

Safe Bayesian Optimization for Complex Control Systems via Additive Gaussian Processes

Model ReleasesDGX agent

arXiv:2408.16307v3 Announce Type: replace-cross Abstract: Automatic controller tuning is attractive for robotics and mechatronic systems whose dynamics are difficult to model accurately, but direct bl

Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image

Model ReleasesDGX agent

arXiv:2605.14984v1 Announce Type: cross Abstract: Generating a street-level 3D scene from a single satellite image is a crucial yet challenging task. Current methods present a stark trade-off: geometr

SD-ReID: View-aware Stable Diffusion for Aerial-Ground Person Re-Identification

ResearchDGX agent

arXiv:2504.09549v3 Announce Type: replace Abstract: Aerial-Ground Person Re-IDentification (AG-ReID) aims to retrieve specific persons across cameras with different viewpoints. Previous works focus on

Selective Safety Steering via Value-Filtered Decoding

SafetyDGX agent

arXiv:2605.14746v1 Announce Type: new Abstract: While large language models (LLMs) are trained to align with human values, their generations may still violate safety constraints. A growing line of wor

Silent Collapse in Recursive Learning Systems

AgentsDGX agent

arXiv:2605.14588v1 Announce Type: new Abstract: Recursive learning -- where models are trained on data generated by previous versions of themselves -- is increasingly common in large language models,

Stochastic Attention via Langevin Dynamics on the Modern Hopfield Energy

Model ReleasesDGX agent

arXiv:2603.06875v3 Announce Type: replace Abstract: Attention heads retrieve: given a query, they return a weighted average of stored values. We showed that this computation is one step of gradient de

The Rate-Distortion-Polysemanticity Tradeoff in SAEs

Model ReleasesDGX agent

arXiv:2605.14694v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) that can accurately reconstruct their input (minimizing distortion) by making efficient use of few features (minimizing the r

Towards the Next Frontier of LLMs, Training on Private Data: A Cross-Domain Benchmark for Federated Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.13936v1 Announce Type: cross Abstract: The recent success of large language models (LLMs) has been largely driven by vast public datasets. However, the next frontier for LLM development lie

VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing

Model ReleasesDGX agent

arXiv:2510.05213v2 Announce Type: replace-cross Abstract: Pretrained vision foundation models (VFMs) advance robotic learning via rich visual representations, yet individual VFMs typically excel only

VGGT-Omega

SafetyDGX agent

arXiv:2605.15195v1 Announce Type: new Abstract: Recent feed-forward reconstruction models, such as VGGT, have proven competitive with traditional optimization-based reconstructors while also providing

VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing

Model ReleasesDGX agent

arXiv:2602.07045v2 Announce Type: replace-cross Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have enabled complex reasoning. However, existing remote sensing (RS) benchmar

VMU-Diff: A Coarse-to-fine Multi-source Data Fusion Framework for Precipitation Nowcasting

ResearchDGX agent

arXiv:2605.14597v1 Announce Type: new Abstract: Precipitation nowcasting is a vital spatio-temporal prediction task for meteorological applications but faces challenges due to the chaotic property of

Watch your neighbors: Training statistically accurate chaotic systems with local phase space information

Local AiDGX agent

arXiv:2605.14405v1 Announce Type: new Abstract: Chaotic systems pose fundamental challenges for data-driven dynamics discovery, as small modeling errors lead to exponentially growing trajectory discre

Web Agents Should Adopt the Plan-Then-Execute Paradigm

Model ReleasesDGX agent

arXiv:2605.14290v1 Announce Type: cross Abstract: ReAct has become the default architecture across LLM agents, and many existing web agents follow this paradigm. We argue that it is the wrong default

Why Retrieval-Augmented Generation Fails: A Graph Perspective

ResearchDGX agent

arXiv:2605.14192v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) has become a powerful and widely used approach for improving large language models by grounding generation in ret

XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge Composition

Model ReleasesDGX agent

arXiv:2605.14754v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed for knowledge synthesis, yet their capacity for compositional generalization in scientific knowle

14 May 2026

1/5 Caching carries a deterministic assumption baked in: same input, same output. That breaks down with LLMs, and especially with agents run…

Model ReleasesDGX agent

This post discusses how traditional caching mechanisms assume deterministic behavior (identical inputs producing identical outputs), an assumption that breaks down with large language models and espec

← Previous
1…413414415416417…1059
Next →