AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
All
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,628 results
Model Releases

Small Data, Big Noise: Adversarial Training for Robust Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2606.10610v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) has become essential for adapting foundation models to downstream NLP tasks. However, current PEFT methods often

model-releasesarxiv-cs-cl
10 Jun 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Sources: Microsoft is restricting employees from using Claude Fable 5 because of Anthropic's new 30-day data retention requirements (Tom Warren/The Verge)

DGX agent

Tom Warren / The Verge: Sources: Microsoft is restricting employees from using Claude Fable 5 because of Anthropic's new 30-day data retention requirements — Microsoft's legal teams are evaluating Ant

model-releasestechmeme
10 Jun 2026
Model Releases

SPACE: Source-free Proxy Anchor Concept Erasure for MLLMs

DGX agent

arXiv:2606.09868v1 Announce Type: cross Abstract: As Multimodal Large Language Models (MLLMs) face growing privacy risks and regulatory constraints, machine unlearning (MU) has emerged as a crucial so

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

SPDM: Geometry-Modulated State Space Modeling with Manifold Constraints for Time Series Forecasting

DGX agent

arXiv:2606.09917v1 Announce Type: new Abstract: Multivariate time series forecasting requires capturing the continuously evolving correlation structure among interacting variables. Existing state-spac

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

SpineReport: Automated 3D Quantification and Reporting of Lumbar Spine Degeneration on MRI

DGX agent

arXiv:2606.10021v1 Announce Type: new Abstract: Lumbar spine conditions are a leading cause of disability worldwide, yet reliable quantification of degeneration from MRI remains challenging. In clinic

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

SSR-Merge: Subspace Signal Routing for Training-Free LoRA Merging in Diffusion Models

DGX agent

arXiv:2606.10617v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) merging can efficiently combine diverse generative capabilities from multiple trained LoRAs for a diffusion model. However, e

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios

DGX agent

arXiv:2606.10394v1 Announce Type: new Abstract: Large language models are increasingly used to power personal agents for everyday applications, but evaluating these agents remains a challenge. Existin

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Streaming Knowledge Compilation: Proactive Materiality-Scored Pinning for Time-Evolving LLM Wikis

DGX agent

arXiv:2606.09877v1 Announce Type: cross Abstract: LLM wiki systems compile knowledge into pre-filled KV caches for efficient inference, but assume a static corpus -- an assumption that fails whenever

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

Structure from Reasoning, Numbers from Search: On-Premise Open LLMs as Structural Priors for Coupled MIMO Controller Tuning

DGX agent

arXiv:2606.11015v1 Announce Type: new Abstract: Tuning controllers for strongly coupled multi-input multi-output (MIMO) industrial processes is hard: decentralized classical auto-tuning ignores loop i

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

DGX agent

arXiv:2606.11070v1 Announce Type: cross Abstract: Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However,

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Temporal Context Conditioning for Seasonality-Aware Precipitation Nowcasting of High-Intensity Rainfall

DGX agent

arXiv:2606.09959v1 Announce Type: cross Abstract: Precipitation nowcasting is increasingly being approached with deep learning models that learn directly from recent radar observations. Although such

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Temporal Sheaf Neural Networks with Dynamic Orthogonal Transport

DGX agent

arXiv:2606.10071v1 Announce Type: cross Abstract: We introduce Temporal Sheaf Neural Networks (TSNN), a temporal link prediction framework that equips each node with a time-varying orthogonal frame an

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

TENP: Trapezoidal Expert Neuron Pruning For Mixture-of-Experts

DGX agent

arXiv:2606.09885v1 Announce Type: new Abstract: Mixture-of-Experts large language models (LLMs) scale efficiently through sparse activation, yet their deployment is fundamentally constrained by the la

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

The 1st PortraitCraft Challenge: A CVPR 2026 Workshop Competition on Portrait Composition Understanding and Generation

DGX agent

arXiv:2606.10894v1 Announce Type: new Abstract: This paper presents an overview of the inaugural PortraitCraft Challenge, held as one of the official competitions at CVPR 2026. The challenge focuses o

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

The hyper-scaled NLP bound for maximum-entropy remote sampling

DGX agent

arXiv:2601.20970v3 Announce Type: replace-cross Abstract: The maximum-entropy remote sampling problem (MERSP) is to select a subset of s random variables from a set of n random variables, so as to max

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

The Interlocutor Effect: Why LLMs Leak More Personal Data to Agents Than Humans

DGX agent

arXiv:2606.09844v1 Announce Type: cross Abstract: Large Language Models (LLMs) alter their privacy behavior based on the perceived identity of their interlocutor. While safety mechanisms typically pre

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

The Order Matters: Sequential Fine-Tuning of LLaMA for Coherent Automated Essay Scoring

DGX agent

arXiv:2606.10327v1 Announce Type: new Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.g., lead, claim, evidence, conclusion), yet most approaches treat

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

the paper is long (30 pages), but have your AI read it: https://arxiv.org/abs/2606.10241 here's a simple interactive tutorial on the topic t…

DGX agent

the paper is long (30 pages), but have your AI read it: https://arxiv.org/abs/2606.10241 here's a simple interactive tutorial on the topic that claude made: https://claude.ai/public/artifacts/038db6cf

model-releasesyohei-nakajima--x
10 Jun 2026
Model Releases

The Shibboleth Effect: Auditing the Cross-Lingual Distributional Skew of Large Language Models

DGX agent

arXiv:2606.11082v1 Announce Type: new Abstract: This study investigates cross-lingual distributional skew (the Shibboleth Effect) in frontier large language models (LLMs) subjected to sustained advers

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

“They have been running literally the same play for seven years: scare, hype (evoking media interest), and (eventually) release. Scare, hype…

DGX agent

“They have been running literally the same play for seven years: scare, hype (evoking media interest), and (eventually) release. Scare, hype, release. And repeat. Is it that hard to see?” 👇🏼 - @GaryMa

model-releasesgary-marcus--x
10 Jun 2026
Model Releases

This is a great article on how startups/frontier labs can coexist. Another way to look at this is task complexity - the number of bits of in…

DGX agent

This is a great article on how startups/frontier labs can coexist. Another way to look at this is task complexity - the number of bits of information needed to specify a task such that AI can solve th

model-releasesjerry-liu--x
10 Jun 2026
Model Releases

Trainable Smooth-Rotation Transforms with Learned Channel Scales for LLM Quantization

DGX agent

arXiv:2606.09927v1 Announce Type: cross Abstract: Post-training quantization (PTQ) is one of the most practical ways to reduce the serving cost of Large Language Models (LLMs), but activation quantiza

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Training LLMs to Enforce Multi-Level Instruction Hierarchies via Gravity-Weighted Direct Preference Optimization

DGX agent

arXiv:2606.10860v1 Announce Type: cross Abstract: Production LLMs receive instructions from sources with very different levels of trust, yet attend to every token with uniform architectural privilege.

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

TRAPS: Therapeutic Response Analysis via Pathway-informed Stratification

DGX agent

arXiv:2606.09898v1 Announce Type: new Abstract: Cancer treatment planning requires decisions across multiple clinical dimensions at once. Clinicians must determine whether a patient should receive tar

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

U-TTT: Towards Generalizable PET Image Denoising via Test-Time Training

DGX agent

arXiv:2606.11032v1 Announce Type: new Abstract: Existing deep learning models for Positron Emission Tomography (PET) image denoising often suffer from severe performance degradation under distribution

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

UMI-Bench 1.0: An Open and Reproducible Real-World Benchmark for Tabletop Robotic Manipulation with UMI Data

DGX agent

arXiv:2606.10382v1 Announce Type: new Abstract: Real-robot evaluation is essential for understanding whether learned manipulation policies can operate reliably outside curated demonstrations. This nee

model-releasesarxiv-cs-ro
10 Jun 2026
Model Releases

Uncertainty-Aware Motion Planning for Autonomous Driving in Mixed Traffic Environment

DGX agent

arXiv:2606.09958v1 Announce Type: cross Abstract: In mixed-traffic environments where autonomous and human-driven vehicles may co-exist, motion planning for autonomous vehicles requires anticipating t

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

UXBench: Benchmarking User Experience in AI Assistants

DGX agent

arXiv:2606.09570v2 Announce Type: replace Abstract: As AI assistants serve millions of users daily, evaluating user experience (UX) beyond general model capability has become increasingly important. W

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions

DGX agent

arXiv:2512.11995v2 Announce Type: replace-cross Abstract: While many vision-language models (VLMs) are developed to answer well-defined, straightforward questions with highly specified targets, as in

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Validation-Stage Combinatorial Fusion Analysis for Imbalanced Credit-Card Fraud Detection

DGX agent

arXiv:2606.10393v1 Announce Type: new Abstract: Credit-card fraud detection is difficult because fraudulent transactions are rare, costly, and unevenly distributed. Strong gradient-boosted tree models

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

Want to learn more about the Far-Field ASR benchmark? Join Treble's webinar tomorrow, June 11th, with @shinjiw_at_cmu, Cohere's @Julianfmack…

DGX agent

Want to learn more about the Far-Field ASR benchmark? Join Treble's webinar tomorrow, June 11th, with @shinjiw_at_cmu, Cohere's @Julianfmack, and other industry leaders discussing the future of far-fi

model-releasescohere--x
10 Jun 2026
Model Releases

We need more real time data on how AI may be impacting the economy - this is a really useful addition.

DGX agent

We need more real time data on how AI may be impacting the economy - this is a really useful addition. Today, the Stanford @DigEconLab launches the AI Economic Indicators, a new platform for tracking

model-releasesethan-mollick--x
10 Jun 2026
Model Releases

WebChallenger: A Reliable and Efficient Generalist Web Agent

DGX agent

arXiv:2606.10423v1 Announce Type: new Abstract: Autonomous web navigation remains challenging for LLM agents, and the strongest generalist systems rely on proprietary reasoning models whose inference

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

What Demonstration Curation Metrics Do to Your Policy

DGX agent

arXiv:2606.10229v1 Announce Type: cross Abstract: We study whether demonstration-curation metrics that detect defective training episodes also improve the downstream behavior-cloning policy that train

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

What Fits (Into Few Tokens) Doesn't Overfit: Compression and Generalization in ML Research Agents

DGX agent

arXiv:2606.11045v1 Announce Type: new Abstract: Reusing a held-out benchmark adaptively should, in principle, invite overfitting. Yet benchmark-driven machine learning (ML) has produced surprisingly l

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

What makes a harness a harness: necessary and sufficient conditions for an agent harness

DGX agent

arXiv:2606.10106v1 Announce Type: cross Abstract: The term agent harness now circulates widely in software engineering with generative artificial intelligence. It names the layer that wraps a language

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents

DGX agent

arXiv:2606.10267v1 Announce Type: cross Abstract: Hierarchical vision-language-action (Hi-VLA) systems have emerged as a promising paradigm for complex robot manipulation, by using high-level VLM plan

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

When Claude Fable kicks off a workflow, the tokens can go very quickly (these aren't Fable tokens, obviously)

DGX agent

Claude Fable efficiently processes tokens at high speed when initiating workflows, demonstrating rapid token consumption during execution. This observation from Ethan Mollick highlights the computatio

model-releasesethan-mollick--x
10 Jun 2026
Model Releases

When Design Rules Break: Benchmark Composition Determines Whether Label Informativeness Predicts GNN Aggregator Choice

DGX agent

arXiv:2606.10249v1 Announce Type: new Abstract: We examine whether graph neural network (GNN) design rules generalize across benchmark families by studying aggregator selection (sum, mean, max) on 24

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

When Do Autoregressive Sequence Models Forecast Physical Wavefields? A Controlled Study on Synthetic Seismograms

DGX agent

arXiv:2606.10868v1 Announce Type: new Abstract: Long-horizon autoregressive forecasting of oscillatory physical signals, such as seismograms, gravitational-wave strain, and similar wavefields is limit

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff

DGX agent

arXiv:2606.09932v1 Announce Type: cross Abstract: Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) has become a standard pipeline for Large Language Model (LLM) post-training. SFT

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Who Brought Easter Eggs to Eid? Auditing Cultural Translation of Math Word Problems Across Diverse Languages and Regions

DGX agent

arXiv:2606.11009v1 Announce Type: new Abstract: Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether thos

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

WHU-Infra3D: A Full-stack Multi-modal Dataset and Benchmark for 3D Roadside Infrastructure Inventory

DGX agent

arXiv:2606.09882v1 Announce Type: new Abstract: The paradigm of digital twin cities is shifting from coarse visual mapping toward more precise and actionable digitization of urban assets. However, exi

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

With SoftBank apparently struggling to get a margin loan against it’s OpenAI shares, it’s maybe time to repost this, from two years ago. The…

DGX agent

With SoftBank apparently struggling to get a margin loan against it’s OpenAI shares, it’s maybe time to repost this, from two years ago. The vast majority of my earlier worries remain: 9 reasons that

model-releasesgary-marcus--x
10 Jun 2026
Model Releases

wooh https://x.com/shadcn/status/2064671802509410806?s=46

DGX agent

wooh https://x.com/shadcn/status/2064671802509410806?s=46 You have Claude Fable for only a few days. Here's how to make the most of it. Introducing /improve: use your most capable model to audit your

model-releasesswyx--x
10 Jun 2026
Model Releases

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

DGX agent

arXiv:2606.11042v1 Announce Type: new Abstract: Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

WorldOlympiad: Can Your World Model Survive a Triathlon?

DGX agent

arXiv:2606.11129v1 Announce Type: new Abstract: We introduce WorldOlympiad, a benchmark for diagnosing video-based world models across physical faithfulness, geometric consistency, and interaction fid

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

Wrote up my initial impressions of Claude Fable 5 - it has a big model smell: slow, expensive and capable of crunching through pretty much e…

DGX agent

Wrote up my initial impressions of Claude Fable 5 - it has a big model smell: slow, expensive and capable of crunching through pretty much everything I threw at it https://simonwillison.net/2026/Jun/9

model-releasessimon-willison--x
10 Jun 2026
← Previous
1…192193194195196…472
Next →