AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,577 results
30 Jun 2026

Reachability Guarantees for Cart-Pole Swing-Up and Stabilization

Model ReleasesDGX agent

arXiv:2606.28627v1 Announce Type: cross Abstract: The cart-pole swing-up is a canonical benchmark for nonlinear control of underactuated systems, yet an end-to-end guarantee linking the global swing-u

Read the full Aston Martin F1 team interview with @aidangomez here: https://www.astonmartinf1.com/en-GB/news/feature/perspectives-aidan-gome…

Model ReleasesDGX agent

This post links to a full interview with Aidan Gomez conducted by the Aston Martin F1 team, published on their official website under their 'Perspectives' feature section. The interview likely covers

Recursive Self-Evolving Agents via Held-Out Selection

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.28374v1 Announce Type: new Abstract: LLM agents are increasingly improved without weight updates by evolving a natural-language artifact, such as reflections, workflows, playbooks, cheatshe

Redefining Maritime Anomaly Detection via Equation-Grounded Synthetic Anomalies

Model ReleasesDGX agent

arXiv:2606.29721v1 Announce Type: cross Abstract: Maritime anomaly detection is essential for ensuring maritime safety, security, and efficient traffic management at sea, with Automatic Identification

RefAlign: Representation Alignment for Reference-to-Video Generation

Model ReleasesDGX agent

arXiv:2603.25743v2 Announce Type: replace Abstract: Reference-to-video (R2V) generation is a controllable video synthesis paradigm that constrains the generation process using both text prompts and re

Rehearsed Multi-Agent Live Product Demonstrations with Real-Time Voice Question Answering

Model ReleasesDGX agent

arXiv:2606.30294v1 Announce Type: new Abstract: Live product demonstrations are a recurring, high-cost activity in software organizations: a human presenter must select features, dispatch the correspo

Reliability-Prioritized Fine-Grained Generation in Multimodal Large

Model ReleasesDGX agent

arXiv:2606.29573v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly expected to generate fine-grained descriptions of visual content. However, we observe and theo

REPAIR-Bench: A Benchmark for Robot Error Perception And Interaction Recovery

Model ReleasesDGX agent

arXiv:2606.29937v1 Announce Type: new Abstract: Understanding how users perceive and respond to robot failures is essential for building robust and trustworthy robot systems. Prior work, however, (i)

Reported Confidence in LLMs Tracks Commitment More Than Correctness

Model ReleasesDGX agent

arXiv:2606.29490v1 Announce Type: cross Abstract: Confidence is an estimate of the probability that a chosen answer is correct. Verbal confidence reports are widely used as uncertainty measures in lar

Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language Models

Model ReleasesDGX agent

arXiv:2606.29196v1 Announce Type: cross Abstract: Do language models know when they are being tested? This question matters for AI safety: a model that recognises an evaluation context could alter its

Research Entity Extraction and Topic Detection from UKRI Grant Proposals

Model ReleasesDGX agent

arXiv:2606.30304v1 Announce Type: cross Abstract: This paper presents preliminary findings from a UKRI-funded Metascience project comparing three LLM-based approaches, GPT-4o, Mistral, and a bespoke a

Residual-Guided Dictionary Learning for Spectrally Accurate Koopman Approximation

Model ReleasesDGX agent

arXiv:2606.29083v1 Announce Type: cross Abstract: Koopman theory promises linear structure in nonlinear dynamics, but numerical Koopman spectra are easy to compute and hard to trust. A finite EDMD mat

Rethinking Generative Reconstruction Attacks against Graph Neural Network Models

Model ReleasesDGX agent

arXiv:2606.29748v1 Announce Type: new Abstract: The application of graph data in numerous disciplines raises the need for gathering and analyzing huge volumes of data, some of which is private and sen

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation

Model ReleasesDGX agent

arXiv:2606.28998v1 Announce Type: cross Abstract: Large Language Model (LLM) alignment trains an LLM using preference data to produce outputs that better meet established quality standards. While LLM

RIPA: Sensory-Vector Prompt Injection Attacks on LLM-Controlled ROS 2 Robots

Model ReleasesDGX agent

arXiv:2606.28649v1 Announce Type: cross Abstract: We present RIPA, the first systematic multi-channel empirical study of prompt injection attacks delivered through the sensory pipeline of a ROS 2-base

RiverONE: Generating Knowledge-Intensive VLM by Simulated Quantum Machines

Model ReleasesDGX agent

arXiv:2606.29966v1 Announce Type: cross Abstract: Quantum computing provides a powerful paradigm for representing and transforming high-dimensional information through superposition, entanglement, and

RoAd-RL: A Unified Library and Benchmark for Robust Adversarial Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.29867v1 Announce Type: cross Abstract: Deep Reinforcement Learning (DRL) has achieved significant success in robotics and autonomous systems, yet remains vulnerable to adversarial perturbat

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

Model ReleasesDGX agent

arXiv:2606.28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning. However, evaluating these videos is chal

Room 2016 for those attending @aiDotEngineer 2:25pm. Will also cover Galactica, early Llama reasoning efforts and more - think this is the f…

Model ReleasesDGX agent

This post announces a conference session at @aiDotEngineer scheduled for 2:25pm in Room 2016, covering topics including the Galactica model and early reasoning efforts in Llama models, with additional

RSGPNet: Geometric Prompting for Remote Sensing Open-Vocabulary Semantic Segmentation

Model ReleasesDGX agent

arXiv:2606.28410v1 Announce Type: cross Abstract: Open-vocabulary semantic segmentation (OVSS) enables text-guided segmentation of unseen objects, breaking fixed-class limitations to achieve open-worl

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence

Model ReleasesDGX agent

arXiv:2606.20515v2 Announce Type: replace Abstract: Real-world spatial intelligence requires reasoning over a continuous and evolving 3D world, yet existing VLMs and tool-augmented agents largely rema

SA-Homo: Scale Adaptive Homography Estimation for Scale Variation Scenarios

Model ReleasesDGX agent

arXiv:2606.30408v1 Announce Type: new Abstract: Homography estimation, as one of the fundamental problems in computer vision, remains challenged by scale variation scenarios where image pairs potentia

SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics

Model ReleasesDGX agent

arXiv:2606.29894v1 Announce Type: cross Abstract: As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theore

SADL: What to Ignore? A Benchmark for Subject-Aware Distractor Localization

Model ReleasesDGX agent

arXiv:2606.30393v1 Announce Type: new Abstract: Photographs frequently contain visual distractors besides foregrounds and backgrounds of the intended subject, competing for attention and weakening com

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing

Model ReleasesDGX agent

arXiv:2606.29887v1 Announce Type: new Abstract: In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to application-specific safety policies,

SAKE: Software Architectural Knowledge Evaluation Benchmark for Large Language Models

Model ReleasesDGX agent

arXiv:2606.29520v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used as assistants across the software development lifecycle, yet their ability to reason about software

SatSplat: Geometrically-Accurate Gaussian Splatting for Satellite Imagery

Model ReleasesDGX agent

arXiv:2606.28581v1 Announce Type: new Abstract: High-resolution satellite imagery demands 3D reconstruction methods that deliver both speed and geometric accuracy. Recent adaptations of 3D Gaussian Sp

Scalar Representations of Neural Network Training Dynamics

Model ReleasesDGX agent

arXiv:2606.30384v1 Announce Type: new Abstract: Training in artificial neural networks can be viewed as a trajectory evolving through a high-dimensional loss landscape. However, the large number of tr

ScAle: Attention Head Scaling as a Minimal Adapter for Spatial Reasoning in Vision Language Models

Model ReleasesDGX agent

arXiv:2606.29579v1 Announce Type: cross Abstract: Spatial reasoning remains a persistent challenge for many vision language models (VLMs), and improving it typically requires fine-tuning with substant

Scaling Textual Gradients via Sampling-Based Momentum

Model ReleasesDGX agent

arXiv:2506.00400v4 Announce Type: replace-cross Abstract: LLM-based prompt optimization, which uses LLM-provided ``textual gradients'' (feedback) to refine prompts, has emerged as an effective method

Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

Model ReleasesDGX agent

arXiv:2606.30616v1 Announce Type: new Abstract: We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. We invest

SCARCE: Scalable Cascade Analysis for Rare-event Characterisation via Embeddings

Model ReleasesDGX agent

arXiv:2606.29623v1 Announce Type: new Abstract: Rare events govern the safety profile of modern AI systems, yet their probabilities are extremely difficult to estimate: direct Monte Carlo requires pro

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation

Model ReleasesDGX agent

arXiv:2606.30124v1 Announce Type: new Abstract: While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they still struggle with the rigorous semant

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

Model ReleasesDGX agent

arXiv:2603.29139v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientific visuali

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories

Model ReleasesDGX agent

arXiv:2606.28589v1 Announce Type: new Abstract: Current approaches to enhance Large Language Model (LLM) reasoning, such as Chain-of-Thought and 'Wait' prompts, primarily encourage models to think mor

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages

Model ReleasesDGX agent

arXiv:2606.28715v1 Announce Type: cross Abstract: While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly understood d

Selective Memory Retention for Long-Horizon LLM Agents

Model ReleasesDGX agent

arXiv:2606.29178v1 Announce Type: new Abstract: When does retention matter for memory-augmented LLM agents? We study this with TraceRetain, a lightweight framework for bounded external memory in froze

Self-Supervised Theorem Discovery in a Formal Axiomatic System

Model ReleasesDGX agent

arXiv:2606.28747v1 Announce Type: new Abstract: Recent artificial intelligence (AI) systems have shown remarkable progress in mathematical reasoning. Many existing approaches, including large language

Semantic-Driven Scale and Spatial Selection for Efficient Cross-Modal Alignment in Referring Remote Sensing Image Segmentation

Model ReleasesDGX agent

arXiv:2606.30244v1 Announce Type: new Abstract: Referring Remote Sensing Image Segmentation (RRSIS) seeks to localize and segment the target object or region specified by a natural language expression

SEO Agent. Security Agent. Design in Claude. Replit Desktop. Shopify on Replit. Skills + Custom Instructions. 450+ integrations. Package Fir…

Model ReleasesDGX agent

SEO Agent. Security Agent. Design in Claude. Replit Desktop. Shopify on Replit. Skills + Custom Instructions. 450+ integrations. Package Firewall. And more! ✨ All shipped in June! 🤯 And there's a good

Sequential Hiring of Contingent Workers Through Learning-Based Optimization

Model ReleasesDGX agent

arXiv:2606.18438v2 Announce Type: replace-cross Abstract: In this paper, we study a sequential workforce management problem in a contingent labor setting with uncertainty in both worker production and

Sequential Planning via Anchored Robotic Keypoints

Model ReleasesDGX agent

arXiv:2606.30613v1 Announce Type: new Abstract: We present Sequential Planning via Anchored Robotic Keypoints, SPARK, a training-free neurosymbolic manipulation system that reaches 43.7% on six LIBERO

SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution

Model ReleasesDGX agent

arXiv:2606.29713v1 Announce Type: cross Abstract: Hallucination is the reliability bottleneck for LLM-based agents, and fact attribution verifiers are the last line of defense -- yet today's verifiers

SFBench: The SciFy Scientific Feasibility Benchmark

Model ReleasesDGX agent

arXiv:2606.29630v1 Announce Type: new Abstract: We present SFBench, a benchmark dataset for evaluating systems that assess the feasibility of scientific claims. SFBench includes 197 claims in material

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

Model ReleasesDGX agent

arXiv:2606.30201v1 Announce Type: cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical

Singular Learning and Occam's Razor in Deep Monomial Networks

Model ReleasesDGX agent

arXiv:2606.28464v1 Announce Type: new Abstract: In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model's architecture. These critical poi

SkelEM: Training-Signal Decoupling of Skeleton and Diffusion for Self-supervised Axial Super-Resolution in Volume Microscopy

Model ReleasesDGX agent

arXiv:2606.30012v1 Announce Type: new Abstract: Volume microscopy, including electron and light microscopy, suffers from severe anisotropic resolution due to physical axial sectioning. Existing self-s

Sonnet 5 is here! This is going to support better long-running agents. Previous Sonnet models were unreliable, so it's great to see the impr…

Model ReleasesDGX agent

Sonnet 5 is here! This is going to support better long-running agents. Previous Sonnet models were unreliable, so it's great to see the improved version that can complete agentic tasks more reliably.

Spatial Deconfounder: Interference-Aware Deconfounding for Spatial Causal Inference

Model ReleasesDGX agent

arXiv:2510.08762v2 Announce Type: replace Abstract: Causal inference in spatial domains faces two intertwined challenges: (1) unmeasured spatial factors, such as weather, air pollution, or mobility, t

Spectral Embedding via Chebyshev Bases for Robust DeepONet Approximation

Model ReleasesDGX agent

arXiv:2512.09165v2 Announce Type: replace Abstract: Deep Operator Networks (DeepONets) have emerged as a powerful framework for data-driven operator learning, providing flexible surrogates for nonline

SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows

Model ReleasesDGX agent

arXiv:2606.29955v1 Announce Type: cross Abstract: Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet benchmarks

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models

Model ReleasesDGX agent

arXiv:2606.29815v1 Announce Type: new Abstract: Evaluating code large language models (Code LLMs) requires reliable detection of data leakage, where benchmark performance is artificially inflated by e

Stability and Concentration in Nonlinear Inverse Problems with Block-Structured Parameters: Lipschitz Geometry, Identifiability, and an Application to Gaussian Splatting

Model ReleasesDGX agent

arXiv:2602.09415v2 Announce Type: replace Abstract: We develop an operator-theoretic framework for stability and statistical concentration in nonlinear inverse problems with block-structured parameter

StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley

Model ReleasesDGX agent

arXiv:2507.07445v3 Announce Type: replace Abstract: Autonomous agents navigating human society must master both production activities and social interactions, yet existing benchmarks rarely evaluate t

Start building with Nano Banana 2 Lite and Gemini Omni Flash

Model ReleasesDGX agent

Nano Banana 2 Lite is Google's fastest and most cost-efficient image generation model, generating images in approximately 4 seconds , while Gemini Omni Flash is a multimodal model supporting high-qual

Statistically Indistinguishable, Operationally Distinct: A Formal Barrier for Tabular Foundation Models

Model ReleasesDGX agent

arXiv:2606.29091v1 Announce Type: cross Abstract: Tabular foundation models cannot reason about data produced by running systems without access to the rules that govern them. We make this statement fa

Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging

Model ReleasesDGX agent

arXiv:2512.01461v2 Announce Type: replace-cross Abstract: Model merging has emerged as a promising paradigm for enabling multi-task capabilities without additional training. However, traditional basic

STEMGym: Benchmarking Sequential Decision-Making under Dose Budgets in Autonomous Electron Microscopy

Model ReleasesDGX agent

arXiv:2606.29592v1 Announce Type: new Abstract: A central premise of autonomous scientific imaging is that smarter navigation, whether Bayesian, RL-based, or otherwise adaptive, is the principal lever

StrucTab: A Structured Optimization Framework for Table Parsing

Model ReleasesDGX agent

arXiv:2606.29905v1 Announce Type: new Abstract: Table parsing aims to convert table images into structured, machine-readable representations, a task requiring the joint perception of complex spatial l

SurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical Robotics

Model ReleasesDGX agent

arXiv:2606.29247v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models represent a promising direction for embodied intelligence in surgical robotics. Despite the prevalence of VLA benchm

← Previous
1…121122123124125…377
Next →