AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,115
  • Agents7,313
  • Applications5,228
  • Concepts5
  • Hardware1,762
  • Industry6,105
  • Local Ai4,756
  • Model Releases22,759
  • Research19,333
  • Safety12,889
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,115
  • Agents7,313
  • Applications5,228
  • Concepts5
  • Hardware1,762
  • Industry6,105
  • Local Ai4,756
  • Model Releases22,759
  • Research19,333
  • Safety12,889
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
All
85,115Total entries
1Added by human
85,114Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,767 results
Model Releases

SVHighlights: Towards Extremely Long Sport Video Highlight Detection

DGX agent

arXiv:2606.06926v1 Announce Type: new Abstract: While highlight detection for long-form videos is of great practical importance, most existing methods remain limited to short-form content, largely due

model-releasesarxiv-cs-cv
8 Jun 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

SW-A^2-Bench: Benchmarking Autonomous Software Agent Generation for Agentic Web

DGX agent

arXiv:2604.04226v2 Announce Type: replace-cross Abstract: The Agentic Web is emerging as a paradigm in which autonomous software agents interact with online resources and with each other to accomplish

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

SWE-Explore: Benchmarking How Coding Agents Explore Repositories

DGX agent

arXiv:2606.07297v1 Announce Type: cross Abstract: Repository-level coding benchmarks such as SWE-bench have driven a rapid surge in the capabilities of coding agents. Yet they usually treat coding tas

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

TALAN: Task-Aligned Latent Adaptation Networks for Targeted Post-Training of Large Language Models

DGX agent

arXiv:2606.06902v1 Announce Type: new Abstract: Targeted post-training aims to improve reasoning, math, and code without degrading strengths. Low-rank adapters are efficient but task-global; activatio

model-releasesarxiv-cs-lg
8 Jun 2026
Model Releases

Test-Time Trajectory Optimization for Autonomous Driving

DGX agent

arXiv:2606.07170v1 Announce Type: new Abstract: End-to-end planners for autonomous driving typically generate a set of candidate trajectories, score each one, and return the highest-scoring candidate.

model-releasesarxiv-cs-ro
8 Jun 2026
Model Releases

TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment

DGX agent

arXiv:2606.07451v1 Announce Type: cross Abstract: Vision-language models such as CLIP are highly useful for diverse tasks due to their shared image-text embedding space. Despite this, the image and te

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

Textual Supervision Enhances Geospatial Representations in Vision-Language Models

DGX agent

arXiv:2606.07172v1 Announce Type: cross Abstract: Geospatial understanding is a critical yet underexplored dimension in the development of machine learning systems for tasks such as image geolocation

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

The Fine-Tuning Trap: Evaluating Negative Transfer and the Role of PEFT in Sub-1B Mathematical Reasoning

DGX agent

arXiv:2606.06920v1 Announce Type: cross Abstract: Deploying Small Language Models (SLMs) on edge devices requires efficient fine-tuning strategies that adapt models to new tasks without degrading thei

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

The Geometry of Representational Failures in Vision Language Models

DGX agent

arXiv:2602.07025v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) exhibit puzzling failures in multi-object visual tasks, such as hallucinating non-existent elements or failing t

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

The highest leverage work in AI right now is some of the most boring. (Well boring to others, I kind of love the pain of problem-solving.) E…

DGX agent

The highest leverage work in AI right now is some of the most boring. (Well boring to others, I kind of love the pain of problem-solving.) Everyone and their boss wants to build the cool AI agent... t

model-releasesallie-k--miller--x
8 Jun 2026
Model Releases

The north stars we're working towards at OpenAI all center around the mission: ensure AGI benefits all of humanity. AI should expand human a…

DGX agent

The north stars we're working towards at OpenAI all center around the mission: ensure AGI benefits all of humanity. AI should expand human agency, not make people less consequential to the future. htt

model-releasesopenai--x
8 Jun 2026
Model Releases

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment

DGX agent

arXiv:2606.06667v1 Announce Type: new Abstract: The mechanisms behind LLMs' broad over-generalization beyond training examples remain unclear. Emergent misalignment (EM) offers a striking case study:

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

The Post-GCN Decade Revisited: Curvature-Stratified Evaluation of Relational Learning

DGX agent

arXiv:2606.06397v2 Announce Type: replace Abstract: Current evaluation practices in relational learning rely heavily on flat leaderboards that average performance across heterogeneous datasets, implic

model-releasesarxiv-cs-lg
8 Jun 2026
Model Releases

There’s so much demand for a good small model, look at top downloaded qwen models All < 9b

DGX agent

The post highlights strong market demand for efficient small language models under 9 billion parameters, citing Qwen's top-downloaded models as evidence of user preference for compact, resource-effici

model-releasesclem-delangue--x
8 Jun 2026
Model Releases

Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models

DGX agent

arXiv:2606.07157v1 Announce Type: new Abstract: Many efforts to ensure frontier AI models are safe rely on monitoring their chain-of-thought (CoT) reasoning. If models become able to perform sufficien

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

Think Like a Pilot: Fine-Grained Long-Horizon UAV Navigation

DGX agent

arXiv:2606.06836v1 Announce Type: cross Abstract: Language-guided UAV agents must execute long-horizon semantic instructions while producing smooth, physically feasible continuous flight commands, yet

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

ThinkBooster: A Unified Framework for Seamless Test-Time Scaling of LLM Reasoning

DGX agent

arXiv:2606.06915v1 Announce Type: cross Abstract: Test-time compute (TTC) scaling has emerged as a powerful paradigm for improving large language model (LLM) reasoning by allocating additional compute

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

This is super big I think this is the first useful speculative decoding method deployed on a big quasi frontier model Massive unlock @fi5662…

DGX agent

This is super big I think this is the first useful speculative decoding method deployed on a big quasi frontier model Massive unlock @fi56622380 🚀 1,000+ TOKENS/S ON A 1T MODEL! 🚀 We are thrilled to r

model-releasesjeremy-howard--x
8 Jun 2026
Model Releases

TokaMind: A Multi-Modal Transformer Foundation Model for Tokamak Plasma Dynamics

DGX agent

arXiv:2602.15084v2 Announce Type: replace-cross Abstract: We present TokaMind, to our knowledge the first open-source foundation model for tokamak plasma dynamics, based on a Multi-Modal Transformer (

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

Tokyo this week

DGX agent

This post likely provides weekly updates, events, or recommendations for activities and happenings in Tokyo for the current week. The content probably covers entertainment, dining, cultural events, or

model-releasesboris-cherny--x
8 Jun 2026
Model Releases

Towards Tight Bounds for Streaming Attention

DGX agent

arXiv:2606.07205v1 Announce Type: cross Abstract: The attention mechanism is a cornerstone of modern transformer architectures. However, its expressive power comes at the cost of quadratic runtime and

model-releasesarxiv-cs-lg
8 Jun 2026
Model Releases

Trading Engagement for Sustainability: Carbon-Aware Re-ranking for E-commerce Recommendations

DGX agent

arXiv:2606.04550v1 Announce Type: cross Abstract: E-commerce recommender systems strongly influence which products users consider and purchase, yet sustainability signals such as Product Carbon Footpr

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

Tree-of-Experience: A Structured Experience-Management Solution for Self-Evolving Agents under Low-Repetition and Implicit-Reward Environments

DGX agent

arXiv:2606.06960v1 Announce Type: new Abstract: Experience-based self-evolution is crucial for LLM agents, but existing benchmarks often assume explicit goals, stable task patterns, and clear feedback

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

TSAQA: Time Series Analysis Question And Answering Benchmark

DGX agent

arXiv:2601.23204v2 Announce Type: replace Abstract: Time series data are integral to critical applications across domains such as finance, healthcare, transportation, and environmental science. While

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

Twin: Tuning Learning Rate and Weight Decay of Deep Homogeneous Classifiers without Validation

DGX agent

arXiv:2403.05532v2 Announce Type: replace-cross Abstract: We introduce Tune without Validation (Twin), a simple and effective pipeline for tuning learning rate and weight decay of homogeneous classifi

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

Uncertainty-Aware LLM-Guided Policy Shaping for Sparse-Reward Reinforcement Learning

DGX agent

arXiv:2606.06673v1 Announce Type: new Abstract: Sparse rewards and heterogeneous task sequences remain persistent challenges in Reinforcement Learning (RL), often resulting in slow convergence, weak g

model-releasesarxiv-cs-lg
8 Jun 2026
Model Releases

Uncertainty-Guided Label Rebalancing for CPS Safety Monitoring

DGX agent

arXiv:2603.25670v3 Announce Type: replace Abstract: Safety monitoring is essential for Cyber-Physical Systems (CPSs). However, unsafe events are rare in real-world CPS operations, creating an extreme

model-releasesarxiv-cs-lg
8 Jun 2026
Model Releases

Uniform Stability and Generalization Error of GD and SGD on Fixed-Point Parameters

DGX agent

arXiv:2606.06934v1 Announce Type: new Abstract: We analyze generalization error, uniform stability, and uniform argument stability of gradient descent (GD) and stochastic gradient descent (SGD) over d

model-releasesarxiv-cs-lg
8 Jun 2026
Model Releases

UniSHARP: Universal Sharp Monocular View Synthesis

DGX agent

arXiv:2606.07514v1 Announce Type: new Abstract: In this work, we focus on extending SHARP, the popular photorealistic view synthesis method, for universal monocular rendering across a continuum of cam

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs

DGX agent

arXiv:2606.06622v1 Announce Type: new Abstract: We introduce UnpredictaBench, an evaluation that tests the ability of large language models (LLMs) to capture true underlying distributions. As LLMs are

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

Unsupervised Continual Clustering via Forward-Backward Knowledge Distillation

DGX agent

arXiv:2606.07474v1 Announce Type: new Abstract: Unsupervised Continual Learning (UCL) aims to enable neural networks to learn sequential tasks without labels or access to past data. A major challenge

model-releasesarxiv-cs-lg
8 Jun 2026
Model Releases

UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding

DGX agent

arXiv:2606.07167v1 Announce Type: cross Abstract: Meaningful multilingual evaluation must test models in the target language and educational context. Urdu, spoken by more than 230 million people, lack

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

VeriDrive: Verifiable Counterfactual Supervision for Cost-Efficient Vision-Language Planning

DGX agent

arXiv:2606.07338v1 Announce Type: new Abstract: Vision-language driving models increasingly use reasoning supervision to bridge perception, prediction, and planning, but existing driving rationales ar

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object Segmentation

DGX agent

arXiv:2606.06819v1 Announce Type: new Abstract: Reasoning Video Object Segmentation (RVOS) demands a sophisticated integration of temporal dynamics, spatial details, and linguistic reasoning to achiev

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

we just shipped support for rubrics in deepagents ✅ give your agent a clear definition of what 'done' looks like, and force it to run in a l…

DGX agent

we just shipped support for rubrics in deepagents ✅ give your agent a clear definition of what 'done' looks like, and force it to run in a loop until said goal is complete this is similar to /goal in

model-releasesharrison-chase--x
8 Jun 2026
Model Releases

What Your Posts Reveal: A Benchmark and Agentic Framework for User-Level Privacy Leakage on Social Media

DGX agent

arXiv:2606.06784v1 Announce Type: cross Abstract: Public social media posts can reveal private information through weak cues scattered across text, images, or metadata. Such leakage is often cumulativ

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations

DGX agent

arXiv:2606.07237v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in healthcare for tasks such as clinical question answering, diagnosis support, and report summariz

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

When Recovery Matters: The Blind Spot of Surrogate Privacy in MLLM Editing

DGX agent

arXiv:2606.07171v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) enable flexible instruction-driven image editing, but privacy risks arise when user images expose diverse and u

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

When to Think Deeply: Inhibitory Deliberation for LLM Reasoning

DGX agent

arXiv:2606.06745v1 Announce Type: new Abstract: Reasoning Large Language Models can improve problem-solving performance through deliberative inference, but invoking slow reasoning for every input is c

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

When we first demoed Claude Code internally, it got two reactions on Slack. A year after GA, @_catwu and I sat down to talk about what's cha…

DGX agent

When we first demoed Claude Code internally, it got two reactions on Slack. A year after GA, @_catwu and I sat down to talk about what's changed: why I use auto mode instead of plan mode, how routines

model-releasesboris-cherny--x
8 Jun 2026
Model Releases

Which Anatomy Matters Under Limited Labels? A Data-Efficient Anatomy-Aware Benchmark for Cardiac Pathology Prediction

DGX agent

arXiv:2606.06509v1 Announce Type: cross Abstract: Numerous medical imaging problems must be solved under limited labels and constrained compute, yet it remains unclear whether performance gains are dr

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

Wordle 1,814 4/6 ⬛🟨⬛⬛⬛ ⬛⬛⬛⬛⬛ ⬛🟨🟨⬛⬛ 🟩🟩🟩🟩🟩

DGX agent

This post documents a Wordle game result where the player solved puzzle #1,814 in four attempts, using color-coded emoji tiles (⬛ for incorrect letters, 🟨 for correct letters in wrong positions, 🟩 for

model-releasesanthropic--x
8 Jun 2026
Model Releases

WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

DGX agent

arXiv:2606.06538v1 Announce Type: new Abstract: In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

Xiaomi claims MiMo-V2.5-Pro-UltraSpeed tops 1K tokens/second, a first at the 1T-parameter scale, using a standard 8-GPU commodity node; API trial starts June 9 (Jose Antonio Lanz/Decrypt)

DGX agent

Jose Antonio Lanz / Decrypt: Xiaomi claims MiMo-V2.5-Pro-UltraSpeed tops 1K tokens/second, a first at the 1T-parameter scale, using a standard 8-GPU commodity node; API trial starts June 9 — Most peop

model-releasestechmeme
8 Jun 2026
Model Releases

Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs

DGX agent

arXiv:2601.12359v1 Announce Type: cross Abstract: Prompt injection attacks have become an increasing vulnerability for LLM applications, where adversarial prompts exploit indirect input channels such

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

A point on the board for 2026. Thank you to the team, our partners, and our fans for the ongoing work and continued support. #MonacoGP

DGX agent

Aston Martin F1 team secured a point at the 2026 Monaco Grand Prix, acknowledging contributions from their team members, partners, and supporters. The post expresses gratitude for the collaborative ef

model-releasescohere--x
7 Jun 2026
Model Releases

datasette-agent-edit 0.1a0

DGX agent

Release: datasette-agent-edit 0.1a0 I'm planning several plugins for Datasette Agent which can make edits to existing pieces of text - things like collaborative Markdown editing, updating large SQL qu

model-releasessimon-willison
7 Jun 2026
Model Releases

Falcon 9 launches 21 @Starlink satellites and two Starshield satellites from California

DGX agent

SpaceX's Falcon 9 rocket successfully launched 21 Starlink internet satellites and two Starshield military satellites from a California launch facility. Starlink satellites are part of SpaceX's global

model-releaseselon-musk--x
7 Jun 2026
← Previous
1…208209210211212…475
Next →