AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,595 results
10 Jun 2026

Sigma-Branch: Hierarchical Single-Path Network Reconstruction for Dynamic Inference with Reduced Active Parameters

Model ReleasesDGX agent

arXiv:2606.09924v1 Announce Type: cross Abstract: Deploying deep neural networks on memory-constrained edge accelerators is bottlenecked by per-inference off-chip weight transfer rather than computati

Sim2Schedule: A Simulator-Guided LLM Framework for Autonomous Open-Pit Mine Scheduling

Model ReleasesDGX agent

arXiv:2606.10286v1 Announce Type: new Abstract: Open-pit mine scheduling is a critical process for maximizing economic return under complex geotechnical and operational constraints. While Mixed-Intege

SkillResolve-Bench: Measuring and Resolving Same-Capability Ambiguity in Agent Skill Retrieval


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2606.10388v1 Announce Type: cross Abstract: Agent skill libraries are becoming routable software assets: a retrieved skill can contribute instructions, scripts, resource bindings, and execution

Small Data, Big Noise: Adversarial Training for Robust Parameter-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2606.10610v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) has become essential for adapting foundation models to downstream NLP tasks. However, current PEFT methods often

Sources: Microsoft is restricting employees from using Claude Fable 5 because of Anthropic's new 30-day data retention requirements (Tom Warren/The Verge)

Model ReleasesDGX agent

Tom Warren / The Verge: Sources: Microsoft is restricting employees from using Claude Fable 5 because of Anthropic's new 30-day data retention requirements — Microsoft's legal teams are evaluating Ant

SPACE: Source-free Proxy Anchor Concept Erasure for MLLMs

Model ReleasesDGX agent

arXiv:2606.09868v1 Announce Type: cross Abstract: As Multimodal Large Language Models (MLLMs) face growing privacy risks and regulatory constraints, machine unlearning (MU) has emerged as a crucial so

SPDM: Geometry-Modulated State Space Modeling with Manifold Constraints for Time Series Forecasting

Model ReleasesDGX agent

arXiv:2606.09917v1 Announce Type: new Abstract: Multivariate time series forecasting requires capturing the continuously evolving correlation structure among interacting variables. Existing state-spac

SpineReport: Automated 3D Quantification and Reporting of Lumbar Spine Degeneration on MRI

Model ReleasesDGX agent

arXiv:2606.10021v1 Announce Type: new Abstract: Lumbar spine conditions are a leading cause of disability worldwide, yet reliable quantification of degeneration from MRI remains challenging. In clinic

SSR-Merge: Subspace Signal Routing for Training-Free LoRA Merging in Diffusion Models

Model ReleasesDGX agent

arXiv:2606.10617v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) merging can efficiently combine diverse generative capabilities from multiple trained LoRAs for a diffusion model. However, e

STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios

Model ReleasesDGX agent

arXiv:2606.10394v1 Announce Type: new Abstract: Large language models are increasingly used to power personal agents for everyday applications, but evaluating these agents remains a challenge. Existin

Streaming Knowledge Compilation: Proactive Materiality-Scored Pinning for Time-Evolving LLM Wikis

Model ReleasesDGX agent

arXiv:2606.09877v1 Announce Type: cross Abstract: LLM wiki systems compile knowledge into pre-filled KV caches for efficient inference, but assume a static corpus -- an assumption that fails whenever

Structure from Reasoning, Numbers from Search: On-Premise Open LLMs as Structural Priors for Coupled MIMO Controller Tuning

Model ReleasesDGX agent

arXiv:2606.11015v1 Announce Type: new Abstract: Tuning controllers for strongly coupled multi-input multi-output (MIMO) industrial processes is hard: decentralized classical auto-tuning ignores loop i

T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

Model ReleasesDGX agent

arXiv:2606.11070v1 Announce Type: cross Abstract: Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However,

Temporal Context Conditioning for Seasonality-Aware Precipitation Nowcasting of High-Intensity Rainfall

Model ReleasesDGX agent

arXiv:2606.09959v1 Announce Type: cross Abstract: Precipitation nowcasting is increasingly being approached with deep learning models that learn directly from recent radar observations. Although such

Temporal Sheaf Neural Networks with Dynamic Orthogonal Transport

Model ReleasesDGX agent

arXiv:2606.10071v1 Announce Type: cross Abstract: We introduce Temporal Sheaf Neural Networks (TSNN), a temporal link prediction framework that equips each node with a time-varying orthogonal frame an

TENP: Trapezoidal Expert Neuron Pruning For Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2606.09885v1 Announce Type: new Abstract: Mixture-of-Experts large language models (LLMs) scale efficiently through sparse activation, yet their deployment is fundamentally constrained by the la

The 1st PortraitCraft Challenge: A CVPR 2026 Workshop Competition on Portrait Composition Understanding and Generation

Model ReleasesDGX agent

arXiv:2606.10894v1 Announce Type: new Abstract: This paper presents an overview of the inaugural PortraitCraft Challenge, held as one of the official competitions at CVPR 2026. The challenge focuses o

The hyper-scaled NLP bound for maximum-entropy remote sampling

Model ReleasesDGX agent

arXiv:2601.20970v3 Announce Type: replace-cross Abstract: The maximum-entropy remote sampling problem (MERSP) is to select a subset of s random variables from a set of n random variables, so as to max

The Interlocutor Effect: Why LLMs Leak More Personal Data to Agents Than Humans

Model ReleasesDGX agent

arXiv:2606.09844v1 Announce Type: cross Abstract: Large Language Models (LLMs) alter their privacy behavior based on the perceived identity of their interlocutor. While safety mechanisms typically pre

The Order Matters: Sequential Fine-Tuning of LLaMA for Coherent Automated Essay Scoring

Model ReleasesDGX agent

arXiv:2606.10327v1 Announce Type: new Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.g., lead, claim, evidence, conclusion), yet most approaches treat

the paper is long (30 pages), but have your AI read it: https://arxiv.org/abs/2606.10241 here's a simple interactive tutorial on the topic t…

Model ReleasesDGX agent

the paper is long (30 pages), but have your AI read it: https://arxiv.org/abs/2606.10241 here's a simple interactive tutorial on the topic that claude made: https://claude.ai/public/artifacts/038db6cf

The Shibboleth Effect: Auditing the Cross-Lingual Distributional Skew of Large Language Models

Model ReleasesDGX agent

arXiv:2606.11082v1 Announce Type: new Abstract: This study investigates cross-lingual distributional skew (the Shibboleth Effect) in frontier large language models (LLMs) subjected to sustained advers

“They have been running literally the same play for seven years: scare, hype (evoking media interest), and (eventually) release. Scare, hype…

Model ReleasesDGX agent

“They have been running literally the same play for seven years: scare, hype (evoking media interest), and (eventually) release. Scare, hype, release. And repeat. Is it that hard to see?” 👇🏼 - @GaryMa

This is a great article on how startups/frontier labs can coexist. Another way to look at this is task complexity - the number of bits of in…

Model ReleasesDGX agent

This is a great article on how startups/frontier labs can coexist. Another way to look at this is task complexity - the number of bits of information needed to specify a task such that AI can solve th

Trainable Smooth-Rotation Transforms with Learned Channel Scales for LLM Quantization

Model ReleasesDGX agent

arXiv:2606.09927v1 Announce Type: cross Abstract: Post-training quantization (PTQ) is one of the most practical ways to reduce the serving cost of Large Language Models (LLMs), but activation quantiza

Training LLMs to Enforce Multi-Level Instruction Hierarchies via Gravity-Weighted Direct Preference Optimization

Model ReleasesDGX agent

arXiv:2606.10860v1 Announce Type: cross Abstract: Production LLMs receive instructions from sources with very different levels of trust, yet attend to every token with uniform architectural privilege.

TRAPS: Therapeutic Response Analysis via Pathway-informed Stratification

Model ReleasesDGX agent

arXiv:2606.09898v1 Announce Type: new Abstract: Cancer treatment planning requires decisions across multiple clinical dimensions at once. Clinicians must determine whether a patient should receive tar

U-TTT: Towards Generalizable PET Image Denoising via Test-Time Training

Model ReleasesDGX agent

arXiv:2606.11032v1 Announce Type: new Abstract: Existing deep learning models for Positron Emission Tomography (PET) image denoising often suffer from severe performance degradation under distribution

UMI-Bench 1.0: An Open and Reproducible Real-World Benchmark for Tabletop Robotic Manipulation with UMI Data

Model ReleasesDGX agent

arXiv:2606.10382v1 Announce Type: new Abstract: Real-robot evaluation is essential for understanding whether learned manipulation policies can operate reliably outside curated demonstrations. This nee

Uncertainty-Aware Motion Planning for Autonomous Driving in Mixed Traffic Environment

Model ReleasesDGX agent

arXiv:2606.09958v1 Announce Type: cross Abstract: In mixed-traffic environments where autonomous and human-driven vehicles may co-exist, motion planning for autonomous vehicles requires anticipating t

UXBench: Benchmarking User Experience in AI Assistants

Model ReleasesDGX agent

arXiv:2606.09570v2 Announce Type: replace Abstract: As AI assistants serve millions of users daily, evaluating user experience (UX) beyond general model capability has become increasingly important. W

V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions

Model ReleasesDGX agent

arXiv:2512.11995v2 Announce Type: replace-cross Abstract: While many vision-language models (VLMs) are developed to answer well-defined, straightforward questions with highly specified targets, as in

Validation-Stage Combinatorial Fusion Analysis for Imbalanced Credit-Card Fraud Detection

Model ReleasesDGX agent

arXiv:2606.10393v1 Announce Type: new Abstract: Credit-card fraud detection is difficult because fraudulent transactions are rare, costly, and unevenly distributed. Strong gradient-boosted tree models

Want to learn more about the Far-Field ASR benchmark? Join Treble's webinar tomorrow, June 11th, with @shinjiw_at_cmu, Cohere's @Julianfmack…

Model ReleasesDGX agent

Want to learn more about the Far-Field ASR benchmark? Join Treble's webinar tomorrow, June 11th, with @shinjiw_at_cmu, Cohere's @Julianfmack, and other industry leaders discussing the future of far-fi

We need more real time data on how AI may be impacting the economy - this is a really useful addition.

Model ReleasesDGX agent

We need more real time data on how AI may be impacting the economy - this is a really useful addition. Today, the Stanford @DigEconLab launches the AI Economic Indicators, a new platform for tracking

WebChallenger: A Reliable and Efficient Generalist Web Agent

Model ReleasesDGX agent

arXiv:2606.10423v1 Announce Type: new Abstract: Autonomous web navigation remains challenging for LLM agents, and the strongest generalist systems rely on proprietary reasoning models whose inference

What Demonstration Curation Metrics Do to Your Policy

Model ReleasesDGX agent

arXiv:2606.10229v1 Announce Type: cross Abstract: We study whether demonstration-curation metrics that detect defective training episodes also improve the downstream behavior-cloning policy that train

What Fits (Into Few Tokens) Doesn't Overfit: Compression and Generalization in ML Research Agents

Model ReleasesDGX agent

arXiv:2606.11045v1 Announce Type: new Abstract: Reusing a held-out benchmark adaptively should, in principle, invite overfitting. Yet benchmark-driven machine learning (ML) has produced surprisingly l

What makes a harness a harness: necessary and sufficient conditions for an agent harness

Model ReleasesDGX agent

arXiv:2606.10106v1 Announce Type: cross Abstract: The term agent harness now circulates widely in software engineering with generative artificial intelligence. It names the layer that wraps a language

What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents

Model ReleasesDGX agent

arXiv:2606.10267v1 Announce Type: cross Abstract: Hierarchical vision-language-action (Hi-VLA) systems have emerged as a promising paradigm for complex robot manipulation, by using high-level VLM plan

When Claude Fable kicks off a workflow, the tokens can go very quickly (these aren't Fable tokens, obviously)

Model ReleasesDGX agent

Claude Fable efficiently processes tokens at high speed when initiating workflows, demonstrating rapid token consumption during execution. This observation from Ethan Mollick highlights the computatio

When Design Rules Break: Benchmark Composition Determines Whether Label Informativeness Predicts GNN Aggregator Choice

Model ReleasesDGX agent

arXiv:2606.10249v1 Announce Type: new Abstract: We examine whether graph neural network (GNN) design rules generalize across benchmark families by studying aggregator selection (sum, mean, max) on 24

When Do Autoregressive Sequence Models Forecast Physical Wavefields? A Controlled Study on Synthetic Seismograms

Model ReleasesDGX agent

arXiv:2606.10868v1 Announce Type: new Abstract: Long-horizon autoregressive forecasting of oscillatory physical signals, such as seismograms, gravitational-wave strain, and similar wavefields is limit

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff

Model ReleasesDGX agent

arXiv:2606.09932v1 Announce Type: cross Abstract: Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) has become a standard pipeline for Large Language Model (LLM) post-training. SFT

Who Brought Easter Eggs to Eid? Auditing Cultural Translation of Math Word Problems Across Diverse Languages and Regions

Model ReleasesDGX agent

arXiv:2606.11009v1 Announce Type: new Abstract: Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether thos

WHU-Infra3D: A Full-stack Multi-modal Dataset and Benchmark for 3D Roadside Infrastructure Inventory

Model ReleasesDGX agent

arXiv:2606.09882v1 Announce Type: new Abstract: The paradigm of digital twin cities is shifting from coarse visual mapping toward more precise and actionable digitization of urban assets. However, exi

With SoftBank apparently struggling to get a margin loan against it’s OpenAI shares, it’s maybe time to repost this, from two years ago. The…

Model ReleasesDGX agent

With SoftBank apparently struggling to get a margin loan against it’s OpenAI shares, it’s maybe time to repost this, from two years ago. The vast majority of my earlier worries remain: 9 reasons that

wooh https://x.com/shadcn/status/2064671802509410806?s=46

Model ReleasesDGX agent

wooh https://x.com/shadcn/status/2064671802509410806?s=46 You have Claude Fable for only a few days. Here's how to make the most of it. Introducing /improve: use your most capable model to audit your

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

Model ReleasesDGX agent

arXiv:2606.11042v1 Announce Type: new Abstract: Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely

WorldOlympiad: Can Your World Model Survive a Triathlon?

Model ReleasesDGX agent

arXiv:2606.11129v1 Announce Type: new Abstract: We introduce WorldOlympiad, a benchmark for diagnosing video-based world models across physical faithfulness, geometric consistency, and interaction fid

Wrote up my initial impressions of Claude Fable 5 - it has a big model smell: slow, expensive and capable of crunching through pretty much e…

Model ReleasesDGX agent

Wrote up my initial impressions of Claude Fable 5 - it has a big model smell: slow, expensive and capable of crunching through pretty much everything I threw at it https://simonwillison.net/2026/Jun/9

XtrAIn: Training-Guided Occlusion for Feature Attribution

Model ReleasesDGX agent

arXiv:2606.10877v1 Announce Type: cross Abstract: Occlusion-based attribution methods provide an intuitive way to estimate feature importance by perturbing input features and measuring the resulting c

👇@zerohedge confirms what I said here and in my May 29 Marcus on AI newsletter. Tokenmaxxing RIP.

Model ReleasesDGX agent

👇@zerohedge confirms what I said here and in my May 29 Marcus on AI newsletter. Tokenmaxxing RIP. This morning @zerohedge out with a report on the death of Tokenmaxxing 'Microsoft’s AI Chief added to

9 Jun 2026

A Baseline Study and Benchmark for Few-Shot Open-Set Action Recognition with Feature Residual Discrimination

Model ReleasesDGX agent

arXiv:2603.04125v2 Announce Type: replace Abstract: Few-Shot Action Recognition (FS-AR) has shown promising results but is often limited by a closed-set assumption that fails in real-world open-set sc

A Comparative Study of Student Perspectives on Technical Writing Feedback Quality: Evaluating LLMs, SLMs, and Humans in Computer Science Topics

Model ReleasesDGX agent

arXiv:2601.11541v2 Announce Type: replace-cross Abstract: To address the scalability of feedback in computer science while mitigating the privacy and cost limitations of commercial Large Language Mode

A Comparison of SSL-Based Feature Extractors and Back-End Classifiers for Spoofing Detection: A Multi-Corpus Training and Cross-Linguistic Analysis

Model ReleasesDGX agent

arXiv:2606.08669v1 Announce Type: cross Abstract: Voice biometric systems face growing threats from spoofing attacks, yet the evaluation of detection models remains inconsistent across datasets. To in

A Dataset for Dynamic Human Preferences for Vision Language Models

Model ReleasesDGX agent

arXiv:2606.07653v1 Announce Type: cross Abstract: Given the increased adoption of Vision Language Models (VLMs) in human-interactive settings, it is important that we evaluate how well these models ca

A Framework for Evaluating and Benchmarking Concept Drift Detection Methods

Model ReleasesDGX agent

arXiv:2606.07789v1 Announce Type: new Abstract: Data stream mining is fundamentally challenged by concept drift, where distributional changes can degrade model performance. Despite the proliferation o

A Multi-modal Agentic Co-pilot for Evidence Grounded Computational Pathology

Model ReleasesDGX agent

arXiv:2606.08093v1 Announce Type: new Abstract: Pathology is the cornerstone of modern medicine, where accurate decision-making relies heavily on evidence-based practices. While artificial intelligenc

A retrieval conditioned rebinding circuit for dynamic entity tracking in large language models

Model ReleasesDGX agent

arXiv:2606.08644v1 Announce Type: cross Abstract: To interpret context correctly and retrieve relevant information, large language models must bind entities to their attributes and update these bindin

← Previous
1…152153154155156…377
Next →