AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,585 results
26 Jun 2026

WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation

Model ReleasesDGX agent

arXiv:2606.26443v1 Announce Type: cross Abstract: A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object

We stress tested many frontier AI models for multimodal medical reasoning (including GPT-5, Claude 3.5, Gemini 2.5 Pro). They’re not ready. …

Model ReleasesDGX agent

We stress tested many frontier AI models for multimodal medical reasoning (including GPT-5, Claude 3.5, Gemini 2.5 Pro). They’re not ready. Faulty reasoning, use of inappropriate shortcuts, hallucinat

What Do Deepfake Benchmarks Measure? An Audit Using Frozen Self-Supervised Representations


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2606.26384v1 Announce Type: new Abstract: As deepfake generators approach perceptual indistinguishability, reliable detection becomes critical. Yet, detectors that score well on benchmarks routi

What happened after 2,000 people tried to hack my AI assistant

Model ReleasesDGX agent

What happened after 2,000 people tried to hack my AI assistant Fernando Irarrázaval ran a challenge on hackmyclaw.com to see if anyone could leak secrets held by his OpenClaw test instance by sending

What We are Missing in Multimodal LLM Evaluation?

Model ReleasesDGX agent

arXiv:2606.26348v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) can process diverse inputs, e.g., text, images, audio, and video, and generate textual responses. While their c

When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents

Model ReleasesDGX agent

arXiv:2602.08995v2 Announce Type: replace Abstract: Computer-use agents (CUAs) have made tremendous progress in the past year, yet they still frequently produce misaligned actions that deviate from th

When to Write and When to Suppress: Route-Specialized Dual Adapters for Memory-Assisted Knowledge Editing

Model ReleasesDGX agent

arXiv:2606.14668v3 Announce Type: replace Abstract: Knowledge editing systems must update selected facts while preserving nearby but irrelevant behavior. This paper studies this problem in a memory-as

Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs

Model ReleasesDGX agent

arXiv:2606.26987v1 Announce Type: cross Abstract: Recent work identified emotion vectors in Claude Sonnet 4.5, which are internal representations that encode emotion concepts, causally influence behav

Which one is Claude is pretty obvious. GLM-5.2 is a beast in some ways, but doesn't have the self-reflective persona of Claude, and isn't re…

Model ReleasesDGX agent

Ethan Mollick compares Claude and GLM-5.2 AI models, noting that while GLM-5.2 excels in certain capabilities, Claude distinguishes itself through its self-reflective persona and other characteristics

Wordle 1,832 4/6 ⬛🟨⬛⬛⬛ ⬛🟨🟨⬛⬛ 🟨⬛🟩🟨⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post documents a Wordle game result where the player solved puzzle #1,832 in four attempts, using the color-coded emoji system (gray for incorrect letters, yellow for correct letters in wrong pos

XMSE-Aware Adaptive Empirical Bayes Estimation

Model ReleasesDGX agent

arXiv:2606.26975v1 Announce Type: cross Abstract: Empirical Bayes (EB) estimators can match the first-order asymptotic risk of maximum likelihood (ML) while behaving very differently at second order:

Yeah it's a real mystery why Vance would identify with an unlikable guy whose entire life revolved around resenting an elite he felt perpetu…

Model ReleasesDGX agent

Yeah it's a real mystery why Vance would identify with an unlikable guy whose entire life revolved around resenting an elite he felt perpetually excluded from despite credentials and success. JD Vance

Zero-Shot Size Transfer for Neural ODEs on Sparse Random Graphs: Graphon Limits and Adjoint Convergence

Model ReleasesDGX agent

arXiv:2606.26662v1 Announce Type: cross Abstract: Graph Neural Differential Equations (GNDEs) model continuous-time graph dynamics by parameterizing Neural ODE velocity fields with Graph Neural Networ

25 Jun 2026

1000 Rallies: An Event-Camera Dataset and Real-Time Learned Ball-State Estimation for Robotic Table Tennis

Model ReleasesDGX agent

arXiv:2606.25620v1 Announce Type: cross Abstract: Robotic table tennis has emerged as a compelling benchmark for real-time robotic perception due to its fast ball dynamics and stringent timing require

2K Retrofit: Entropy-Guided Efficient Sparse Refinement for High-Resolution 3D Geometry Prediction

Model ReleasesDGX agent

arXiv:2603.19964v3 Announce Type: replace Abstract: High-resolution geometric prediction is essential for robust perception in autonomous driving, robotics, and AR/MR, but current foundation models ar

A 3D-Printable Dataset for Fair Testing and Comparisons of Tactile Sensors

Model ReleasesDGX agent

arXiv:2606.25886v1 Announce Type: cross Abstract: Existing texture datasets for tactile sensing primarily consist of sensor readings from a specific sensor interacting with available surfaces/objects

A Benchmark for Heterogeneous Stereo Deblurring with Physically- and Epipolar-constrained Cross Attention

Model ReleasesDGX agent

arXiv:2606.25962v1 Announce Type: new Abstract: Modern stereo-capable smartphones enable immersive XR content capture. However, hardware heterogeneity across camera modules often causes severe asymmet

A Flow-rate-conserving CNN-based Domain Decomposition Method for Blood Flow Simulations

Model ReleasesDGX agent

arXiv:2509.15900v2 Announce Type: replace-cross Abstract: This work aims to predict blood flow with non-Newtonian viscosity in stenosed arteries using convolutional neural network (CNN) surrogate mode

A Hybrid CNN-LSTM Intrusion Detection Framework for Cybersecurity in Smart Renewable Energy Grids

Model ReleasesDGX agent

arXiv:2606.25200v1 Announce Type: new Abstract: The accelerated digitalization of renewable energy smart grids through IoT sensors, AMI, and SCADA systems has significantly expanded the attack surface

A Leakage-Aware Comparative Benchmark of Machine Learning, Deep Learning, and Transformer Models for Reliable Leukemia Detection

Model ReleasesDGX agent

arXiv:2606.24944v1 Announce Type: cross Abstract: Automated classification of acute lymphoblastic leukemia (ALL) from peripheral blood smear images has often reported near-perfect performance on the C

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

Model ReleasesDGX agent

arXiv:2606.25476v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable performance across natural language processing tasks, yet their deployment in high-stakes appl

A Single Stepsize Suffices for Unprojected Linear TD(0): Simultaneous Robust and Fast Rates via Polyak--Ruppert Averaging

Model ReleasesDGX agent

arXiv:2606.24981v1 Announce Type: new Abstract: We study linear TD(0) under Markovian sampling, where data are generated along a single trajectory. We provide high-probability guarantees for a plain u

Agent-as-a-Router: Agentic Model Routing for Coding Tasks

Model ReleasesDGX agent

arXiv:2606.22902v2 Announce Type: replace Abstract: Real-world users typically have access to multiple Large Language Models (LLMs) from different providers, and these LLMs often excel at distinct dom

Agentic evolution of physically constrained foundation models

Model ReleasesDGX agent

arXiv:2606.25532v1 Announce Type: cross Abstract: Artificial intelligence increasingly drives automated scientific discovery, yet contemporary generalist agents lack physical grounding, frequently hal

AIChilles: Automatically Uncovering Hidden Weaknesses in AI-Evolved Systems

Model ReleasesDGX agent

arXiv:2606.15834v2 Announce Type: replace Abstract: The computer systems community has recently seen growing interest in AI-driven system evolution, where AI agents iteratively rewrite systems. Framew

AISPO: Enhancing Depth Reliability for Robotic Manipulation of Non-Lambertian Objects via Affine-Invariant Shape Prior

Model ReleasesDGX agent

arXiv:2606.25503v1 Announce Type: cross Abstract: Reliable depth perception is critical for robotic manipulation, especially for non-Lambertian objects such as transparent or highly specular surfaces,

AMVICC: A Novel Benchmark for Cross-Modal Failure Mode Profiling for VLMs and IGMs

Model ReleasesDGX agent

arXiv:2601.17037v2 Announce Type: replace Abstract: We investigate visual reasoning limitations of both multimodal large language models (MLLMs) and image generation models (IGMs) by creating a novel

An iterative energy-based multimodal transformer for joint retrieval of wheat soil moisture, leaf area index, and plant height from Sentinel-1 and Sentinel-2 time series

Model ReleasesDGX agent

arXiv:2606.25174v1 Announce Type: cross Abstract: Field-scale retrieval of surface soil moisture (SM), leaf area index (LAI), and plant height (PH) is essential for precision agriculture, yet it remai

Anthropic says Alibaba must be punished for largest Claude cloning attack

Model ReleasesDGX agent

Anthropic accused Alibaba of a coordinated, industrial-scale operation to illicitly scrape data from Claude , using nearly 25,000 fraudulent accounts to generate more than 28.8 million exchanges betwe

Are Tabular Foundation Models Robust to Realistic Query Distribution Shifts in Microbiome Data?

Model ReleasesDGX agent

arXiv:2606.24995v1 Announce Type: new Abstract: Tabular foundation models (TFMs) achieve strong performance on microbiome abundance data, yet their robustness under realistic distribution shift remain

Are We There Yet? Exploring the Capabilities of MLLMs in Assistive AI Applications

Model ReleasesDGX agent

arXiv:2606.25084v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have redefined visual understanding by combining vision encoders with large-scale language models. This unified

Auto-exploration for online reinforcement learning

Model ReleasesDGX agent

arXiv:2512.06244v2 Announce Type: replace Abstract: The exploration-exploitation dilemma in reinforcement learning (RL) is a fundamental challenge to efficient RL algorithms. Existing algorithms for f

Auto-Labelling-Based Domain Transfer for 3D Object Detection on a Bicycle-Mounted LiDAR Platform

Model ReleasesDGX agent

arXiv:2606.25652v1 Announce Type: new Abstract: Reliable 3D perception of vulnerable road users (VRUs) such as cyclists and pedestrians is essential for their safety in urban traffic and a core requir

Benchmarking Deep Learning Models for Laryngeal Cancer Staging Using the LaryngealCT Dataset

Model ReleasesDGX agent

arXiv:2510.11047v2 Announce Type: replace Abstract: Laryngeal cancer imaging research lacks standardised public datasets to enable reproducible deep learning (DL) model development. We present Larynge

Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability

Model ReleasesDGX agent

arXiv:2606.25819v1 Announce Type: new Abstract: Large language models are increasingly deployed as agents that solve tasks by interacting with external tool environments. Although recent tool-use benc

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors

Model ReleasesDGX agent

arXiv:2606.25527v1 Announce Type: new Abstract: Online reinforcement learning (RL) agents increasingly depend on knowledge acquired offline to achieve practical efficiency. Originally studied in offli

Beyond Visual Forensics: Auditing Multimodal Robustness for Synthetic Medical Image Detection

Model ReleasesDGX agent

arXiv:2606.25375v1 Announce Type: new Abstract: With the rapid adoption of generative AI, synthetic medical images pose growing risks, including diagnostic deception and insurance fraud. Although prio

Big Breaking News: White House asks OpenAI to delay GPT- 5.6.

Model ReleasesDGX agent

Big Breaking News: White House asks OpenAI to delay GPT- 5.6. NEW: Trump admin asks OpenAI to stagger GPT-5.6 release over cyber concerns. Will approve 'access customer by customer during this preview

BOFA: Bridge-Layer Orthogonal Low-Rank Fusion for CLIP-Based Class-Incremental Learning

Model ReleasesDGX agent

arXiv:2511.11421v2 Announce Type: replace Abstract: Class-Incremental Learning (CIL) aims to continually learn new categories without forgetting previously acquired knowledge. Vision-language models s

BrainAgent: A Large Language Model-Driven Multi-Agent Framework for Autonomous Brain Signal Understanding

Model ReleasesDGX agent

arXiv:2606.25400v1 Announce Type: new Abstract: Brain-Computer Interfaces (BCIs) and brain signal understanding are pivotal for clinical health and next-generation interactions. Despite this significa

BreachRx launches Rex Platform to coordinate AI-era incident response

Model ReleasesDGX agent

Incident response company BreachRx Inc. today launched the Rex Platform, an agentic artificial intelligence incident command center built for a future in which AI-accelerated attacks set off several b

Building Agent Skills is great, but testing them clearly and across agents is difficult. The Pinecone DevRel team has released Cultivar, a C…

Model ReleasesDGX agent

Building Agent Skills is great, but testing them clearly and across agents is difficult. The Pinecone DevRel team has released Cultivar, a CLI tool and agent skill designed to solve this problem. With

C3-Bench: A Context-Aware Change Captioning Benchmark

Model ReleasesDGX agent

arXiv:2606.25445v1 Announce Type: new Abstract: While Change Captioning systems have garnered substantial attention to respond to our evolving world, their true performance on diverse real-world chang

California launches a tool to serve as an 'early warning system' for widespread AI-driven job loss, linking AI exposure with unemployment insurance claims (Jo Constantz/Bloomberg)

Model ReleasesDGX agent

Jo Constantz / Bloomberg: California launches a tool to serve as an “early warning system” for widespread AI-driven job loss, linking AI exposure with unemployment insurance claims — Politicians like

CausalRAG2: Hierarchical Causal Knowledge Graph Design for RAG

Model ReleasesDGX agent

arXiv:2602.05143v2 Announce Type: replace Abstract: Retrieval augmented generation (RAG) has enhanced large language models by enabling access to external knowledge, with graph-based RAG emerging as a

Claude Code in Slack

Model ReleasesDGX agent

Claude Code is now available in Slack, allowing users to run code directly within the messaging platform. This integration enables developers and teams to execute scripts, test code snippets, and coll

CNN reported in April that US intelligence assessed that roughly half of Iran’s missile launchers had survived US strikes. A more recent IC …

Model ReleasesDGX agent

CNN reported in April that US intelligence assessed that roughly half of Iran’s missile launchers had survived US strikes. A more recent IC report increased that figure to two thirds partially due to

CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks

Model ReleasesDGX agent

arXiv:2604.03314v2 Announce Type: replace-cross Abstract: Foundation models have revolutionized AI, but adapting them efficiently for multimodal tasks, particularly in dual-stream architectures compos

Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos

Model ReleasesDGX agent

arXiv:2603.25645v2 Announce Type: replace-cross Abstract: Early screening via colonoscopy is critical for colon cancer prevention, yet developing robust AI systems for this domain is hindered by the l

Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints

Model ReleasesDGX agent

arXiv:2606.25605v1 Announce Type: new Abstract: Tool Calling and Structured Output are two core capabilities of modern Agent systems, yet their interaction under joint deployment conditions remains in

Curvature-Guided Mixing for MLLM Adaptation

Model ReleasesDGX agent

arXiv:2606.24963v1 Announce Type: new Abstract: Fine-tuning Multimodal Large Language Models (MLLMs) on specialized tasks often leads to catastrophic forgetting of their general capabilities. Existing

Deep Neural Networks with Ordinal Loss for Medical Applications

Model ReleasesDGX agent

arXiv:2606.25769v1 Announce Type: new Abstract: In many prediction problems in medical applications, target labels exhibit an inherent ordinal structure, where class ordering reflects clinically meani

Delta-Position Estimation-Based IMU Odometry: A Comparison of MLP and Kolmogorov-Arnold Networks

Model ReleasesDGX agent

arXiv:2606.25454v1 Announce Type: new Abstract: In this study, the learning-based inertial odometry problem is investigated using raw IMU measurements obtained from the EuRoC MAV benchmark dataset. In

Detect, Unlearn, Restore: Defending Text Summarization Models Against Data Poisoning

Model ReleasesDGX agent

arXiv:2606.26036v1 Announce Type: new Abstract: Training-time data poisoning during fine-tuning poses a significant threat to large language models (LLMs) deployed for abstractive text summarization,

Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography

Model ReleasesDGX agent

arXiv:2606.25546v1 Announce Type: new Abstract: Vision-language pre-training (VLP) holds great promise for general-purpose medical AI by leveraging radiology reports as rich textual supervision, yet e

Distill on a Diet: Efficient Knowledge Distillation via Learnable Data Pruning

Model ReleasesDGX agent

arXiv:2606.25488v1 Announce Type: new Abstract: Knowledge Distillation (KD) is widely used to obtain compact models for efficient inference in resource-constrained environments. Yet the computational

Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning

Model ReleasesDGX agent

arXiv:2510.04773v2 Announce Type: replace Abstract: As Large Language Models (LLMs) demonstrate remarkable capabilities learned from vast corpora, concerns regarding data privacy and safety are receiv

Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation

Model ReleasesDGX agent

arXiv:2606.25782v1 Announce Type: new Abstract: With the widespread adoption of large language models (LLMs) in chatbots and everyday applications, companies increasingly need guardrails that are effe

Do Thinking Tokens Help with Safety?

Model ReleasesDGX agent

arXiv:2606.25013v1 Announce Type: cross Abstract: Today's reasoning models use thinking tokens to attain stronger performance on benchmarks than their instruction-tuned counterparts. It is also genera

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms

Model ReleasesDGX agent

arXiv:2606.25066v1 Announce Type: cross Abstract: Visual search has been one of the most productive paradigms in the study of visual attention: the way reaction time scales with the number of items di

← Previous
1…131132133134135…377
Next →