AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,574 results
30 Jun 2026

RiverONE: Generating Knowledge-Intensive VLM by Simulated Quantum Machines

Model ReleasesDGX agent

arXiv:2606.29966v1 Announce Type: cross Abstract: Quantum computing provides a powerful paradigm for representing and transforming high-dimensional information through superposition, entanglement, and

RoAd-RL: A Unified Library and Benchmark for Robust Adversarial Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.29867v1 Announce Type: cross Abstract: Deep Reinforcement Learning (DRL) has achieved significant success in robotics and autonomous systems, yet remains vulnerable to adversarial perturbat

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning. However, evaluating these videos is chal

Room 2016 for those attending @aiDotEngineer 2:25pm. Will also cover Galactica, early Llama reasoning efforts and more - think this is the f…

Model ReleasesDGX agent

This post announces a conference session at @aiDotEngineer scheduled for 2:25pm in Room 2016, covering topics including the Galactica model and early reasoning efforts in Llama models, with additional

RSGPNet: Geometric Prompting for Remote Sensing Open-Vocabulary Semantic Segmentation

Model ReleasesDGX agent

arXiv:2606.28410v1 Announce Type: cross Abstract: Open-vocabulary semantic segmentation (OVSS) enables text-guided segmentation of unseen objects, breaking fixed-class limitations to achieve open-worl

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence

Model ReleasesDGX agent

arXiv:2606.20515v2 Announce Type: replace Abstract: Real-world spatial intelligence requires reasoning over a continuous and evolving 3D world, yet existing VLMs and tool-augmented agents largely rema

SA-Homo: Scale Adaptive Homography Estimation for Scale Variation Scenarios

Model ReleasesDGX agent

arXiv:2606.30408v1 Announce Type: new Abstract: Homography estimation, as one of the fundamental problems in computer vision, remains challenged by scale variation scenarios where image pairs potentia

SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics

Model ReleasesDGX agent

arXiv:2606.29894v1 Announce Type: cross Abstract: As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theore

SADL: What to Ignore? A Benchmark for Subject-Aware Distractor Localization

Model ReleasesDGX agent

arXiv:2606.30393v1 Announce Type: new Abstract: Photographs frequently contain visual distractors besides foregrounds and backgrounds of the intended subject, competing for attention and weakening com

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing

Model ReleasesDGX agent

arXiv:2606.29887v1 Announce Type: new Abstract: In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to application-specific safety policies,

SAKE: Software Architectural Knowledge Evaluation Benchmark for Large Language Models

Model ReleasesDGX agent

arXiv:2606.29520v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used as assistants across the software development lifecycle, yet their ability to reason about software

SatSplat: Geometrically-Accurate Gaussian Splatting for Satellite Imagery

Model ReleasesDGX agent

arXiv:2606.28581v1 Announce Type: new Abstract: High-resolution satellite imagery demands 3D reconstruction methods that deliver both speed and geometric accuracy. Recent adaptations of 3D Gaussian Sp

Scalar Representations of Neural Network Training Dynamics

Model ReleasesDGX agent

arXiv:2606.30384v1 Announce Type: new Abstract: Training in artificial neural networks can be viewed as a trajectory evolving through a high-dimensional loss landscape. However, the large number of tr

ScAle: Attention Head Scaling as a Minimal Adapter for Spatial Reasoning in Vision Language Models

Model ReleasesDGX agent

arXiv:2606.29579v1 Announce Type: cross Abstract: Spatial reasoning remains a persistent challenge for many vision language models (VLMs), and improving it typically requires fine-tuning with substant

Scaling Textual Gradients via Sampling-Based Momentum

Model ReleasesDGX agent

arXiv:2506.00400v4 Announce Type: replace-cross Abstract: LLM-based prompt optimization, which uses LLM-provided ``textual gradients'' (feedback) to refine prompts, has emerged as an effective method

Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

Model ReleasesDGX agent

arXiv:2606.30616v1 Announce Type: new Abstract: We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. We invest

SCARCE: Scalable Cascade Analysis for Rare-event Characterisation via Embeddings

Model ReleasesDGX agent

arXiv:2606.29623v1 Announce Type: new Abstract: Rare events govern the safety profile of modern AI systems, yet their probabilities are extremely difficult to estimate: direct Monte Carlo requires pro

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation

Model ReleasesDGX agent

arXiv:2606.30124v1 Announce Type: new Abstract: While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they still struggle with the rigorous semant

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

Model ReleasesDGX agent

arXiv:2603.29139v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientific visuali

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories

Model ReleasesDGX agent

arXiv:2606.28589v1 Announce Type: new Abstract: Current approaches to enhance Large Language Model (LLM) reasoning, such as Chain-of-Thought and 'Wait' prompts, primarily encourage models to think mor

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages

Model ReleasesDGX agent

arXiv:2606.28715v1 Announce Type: cross Abstract: While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly understood d

Selective Memory Retention for Long-Horizon LLM Agents

Model ReleasesDGX agent

arXiv:2606.29178v1 Announce Type: new Abstract: When does retention matter for memory-augmented LLM agents? We study this with TraceRetain, a lightweight framework for bounded external memory in froze

Self-Supervised Theorem Discovery in a Formal Axiomatic System

Model ReleasesDGX agent

arXiv:2606.28747v1 Announce Type: new Abstract: Recent artificial intelligence (AI) systems have shown remarkable progress in mathematical reasoning. Many existing approaches, including large language

Semantic-Driven Scale and Spatial Selection for Efficient Cross-Modal Alignment in Referring Remote Sensing Image Segmentation

Model ReleasesDGX agent

arXiv:2606.30244v1 Announce Type: new Abstract: Referring Remote Sensing Image Segmentation (RRSIS) seeks to localize and segment the target object or region specified by a natural language expression

SEO Agent. Security Agent. Design in Claude. Replit Desktop. Shopify on Replit. Skills + Custom Instructions. 450+ integrations. Package Fir…

Model ReleasesDGX agent

SEO Agent. Security Agent. Design in Claude. Replit Desktop. Shopify on Replit. Skills + Custom Instructions. 450+ integrations. Package Firewall. And more! ✨ All shipped in June! 🤯 And there's a good

Sequential Hiring of Contingent Workers Through Learning-Based Optimization

Model ReleasesDGX agent

arXiv:2606.18438v2 Announce Type: replace-cross Abstract: In this paper, we study a sequential workforce management problem in a contingent labor setting with uncertainty in both worker production and

Sequential Planning via Anchored Robotic Keypoints

Model ReleasesDGX agent

arXiv:2606.30613v1 Announce Type: new Abstract: We present Sequential Planning via Anchored Robotic Keypoints, SPARK, a training-free neurosymbolic manipulation system that reaches 43.7% on six LIBERO

SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution

Model ReleasesDGX agent

arXiv:2606.29713v1 Announce Type: cross Abstract: Hallucination is the reliability bottleneck for LLM-based agents, and fact attribution verifiers are the last line of defense -- yet today's verifiers

SFBench: The SciFy Scientific Feasibility Benchmark

Model ReleasesDGX agent

arXiv:2606.29630v1 Announce Type: new Abstract: We present SFBench, a benchmark dataset for evaluating systems that assess the feasibility of scientific claims. SFBench includes 197 claims in material

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

Model ReleasesDGX agent

arXiv:2606.30201v1 Announce Type: cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical

Singular Learning and Occam's Razor in Deep Monomial Networks

Model ReleasesDGX agent

arXiv:2606.28464v1 Announce Type: new Abstract: In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model's architecture. These critical poi

SkelEM: Training-Signal Decoupling of Skeleton and Diffusion for Self-supervised Axial Super-Resolution in Volume Microscopy

Model ReleasesDGX agent

arXiv:2606.30012v1 Announce Type: new Abstract: Volume microscopy, including electron and light microscopy, suffers from severe anisotropic resolution due to physical axial sectioning. Existing self-s

Sonnet 5 is here! This is going to support better long-running agents. Previous Sonnet models were unreliable, so it's great to see the impr…

Model ReleasesDGX agent

Sonnet 5 is here! This is going to support better long-running agents. Previous Sonnet models were unreliable, so it's great to see the improved version that can complete agentic tasks more reliably.

Spatial Deconfounder: Interference-Aware Deconfounding for Spatial Causal Inference

Model ReleasesDGX agent

arXiv:2510.08762v2 Announce Type: replace Abstract: Causal inference in spatial domains faces two intertwined challenges: (1) unmeasured spatial factors, such as weather, air pollution, or mobility, t

Spectral Embedding via Chebyshev Bases for Robust DeepONet Approximation

Model ReleasesDGX agent

arXiv:2512.09165v2 Announce Type: replace Abstract: Deep Operator Networks (DeepONets) have emerged as a powerful framework for data-driven operator learning, providing flexible surrogates for nonline

SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows

Model ReleasesDGX agent

arXiv:2606.29955v1 Announce Type: cross Abstract: Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet benchmarks

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models

Model ReleasesDGX agent

arXiv:2606.29815v1 Announce Type: new Abstract: Evaluating code large language models (Code LLMs) requires reliable detection of data leakage, where benchmark performance is artificially inflated by e

Stability and Concentration in Nonlinear Inverse Problems with Block-Structured Parameters: Lipschitz Geometry, Identifiability, and an Application to Gaussian Splatting

Model ReleasesDGX agent

arXiv:2602.09415v2 Announce Type: replace Abstract: We develop an operator-theoretic framework for stability and statistical concentration in nonlinear inverse problems with block-structured parameter

StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley

Model ReleasesDGX agent

arXiv:2507.07445v3 Announce Type: replace Abstract: Autonomous agents navigating human society must master both production activities and social interactions, yet existing benchmarks rarely evaluate t

Start building with Nano Banana 2 Lite and Gemini Omni Flash

Model ReleasesDGX agent

Nano Banana 2 Lite is Google's fastest and most cost-efficient image generation model, generating images in approximately 4 seconds , while Gemini Omni Flash is a multimodal model supporting high-qual

Statistically Indistinguishable, Operationally Distinct: A Formal Barrier for Tabular Foundation Models

Model ReleasesDGX agent

arXiv:2606.29091v1 Announce Type: cross Abstract: Tabular foundation models cannot reason about data produced by running systems without access to the rules that govern them. We make this statement fa

Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging

Model ReleasesDGX agent

arXiv:2512.01461v2 Announce Type: replace-cross Abstract: Model merging has emerged as a promising paradigm for enabling multi-task capabilities without additional training. However, traditional basic

STEMGym: Benchmarking Sequential Decision-Making under Dose Budgets in Autonomous Electron Microscopy

Model ReleasesDGX agent

arXiv:2606.29592v1 Announce Type: new Abstract: A central premise of autonomous scientific imaging is that smarter navigation, whether Bayesian, RL-based, or otherwise adaptive, is the principal lever

StrucTab: A Structured Optimization Framework for Table Parsing

Model ReleasesDGX agent

arXiv:2606.29905v1 Announce Type: new Abstract: Table parsing aims to convert table images into structured, machine-readable representations, a task requiring the joint perception of complex spatial l

SurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical Robotics

Model ReleasesDGX agent

arXiv:2606.29247v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models represent a promising direction for embodied intelligence in surgical robotics. Despite the prevalence of VLA benchm

SVC-Probe: A Framework for Evaluating Perturbation Generalization in Spatial Foundation-Model Embeddings

Model ReleasesDGX agent

arXiv:2606.28465v1 Announce Type: cross Abstract: This work examines perturbation generalization in spatial foundation-model embeddings derived from fluorescence microscopy images. Although these mode

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance

Model ReleasesDGX agent

arXiv:2603.12703v3 Announce Type: replace Abstract: Video understanding requires models to continuously track and update world state during playback. Although existing benchmarks have advanced video u

SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?

Model ReleasesDGX agent

arXiv:2511.06090v3 Announce Type: replace-cross Abstract: Optimizing the performance of large-scale software repositories demands expertise in code reasoning and software engineering (SWE) to reduce r

SWE-Together: Evaluating Coding Agents in Interactive User Sessions

Model ReleasesDGX agent

arXiv:2606.29957v1 Announce Type: cross Abstract: Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code. Real coding assi

SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios

Model ReleasesDGX agent

arXiv:2511.17649v4 Announce Type: replace-cross Abstract: Tangible control interfaces (TCIs), such as appliance panels, remotes, elevators, and embedded GUIs, are a fundamental component of everyday h

Symbolic Mechanistic Data Attribution: Tracing Training Influence to Learned Behavioral Policies

Model ReleasesDGX agent

arXiv:2606.29171v1 Announce Type: cross Abstract: While existing data attribution methods can identify which training examples build specific mechanistic circuits, they cannot explain how training dat

TextClusterLab: An Integrated Framework for Reliable Text Clustering Studies

Model ReleasesDGX agent

arXiv:2606.28328v1 Announce Type: cross Abstract: In recent years, text clustering has become a critical technique for applications including intent discovery, topic mining, and recommendation systems

TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation

Model ReleasesDGX agent

arXiv:2606.29575v1 Announce Type: cross Abstract: Recent advances in speech separation (SS) have led to compact front-end models with small parameter sizes, yet their high computational cost remains a

Thank you to everyone to came to the Claude managed agents workshop at @aiDotEngineer with @gcemaj and I. We had an absolute blast sharing o…

Model ReleasesDGX agent

Thank you to everyone to came to the Claude managed agents workshop at @aiDotEngineer with @gcemaj and I. We had an absolute blast sharing our journey and walking you through building your first agent

The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling

Model ReleasesDGX agent

arXiv:2606.29278v1 Announce Type: new Abstract: We introduce the Complexity Ceiling Benchmark (CCB), a controlled evaluation of how language-model reasoning decays as the number of required sequential

The Contagion Tensor: A Framework for Measuring Output-Distribution Coupling in Multi-Agent LLM Systems -- and Auditing the Claims It Enables

Model ReleasesDGX agent

arXiv:2606.28839v1 Announce Type: new Abstract: We introduce the Contagion Tensor, a measurement framework for quantifying how large language model (LLM) output distributions couple across modalities,

The CRISTAL Method: Neurosymbolic analysis from AI-synthesized world models

Model ReleasesDGX agent

arXiv:2606.29799v1 Announce Type: new Abstract: This project introduces the CRISTAL Method (Coherent Reliable Intentional Synthesis of Truthful Analysis Logic), a neurosymbolic framework for automatin

The Digital Afterlife of Empires: Four Language Models Converge on the Same Imperial Cartography of Writing

Model ReleasesDGX agent

arXiv:2606.28325v1 Announce Type: cross Abstract: Large language models process the world's writing systems with radical inequality. We constructed the Digital Script Representation Index (DSRI), a se

The FIL Hypothesis: Inductive Biases Help with Kernel Engineering

Model ReleasesDGX agent

arXiv:2606.30442v1 Announce Type: new Abstract: The Bitter Lesson, which posits that general-purpose methods that scale with computation and data ultimately outperform those with built-in human knowle

The food delivery leader in China just dropped an open weights 1.6 trillion parameter model while half the USA was sleeping … 🫨 🇨🇳

Model ReleasesDGX agent

The food delivery leader in China just dropped an open weights 1.6 trillion parameter model while half the USA was sleeping … 🫨 🇨🇳 Meituan, China's largest food delivery platform, open-sourced a 1.6 t

← Previous
1…122123124125126…377
Next →