AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

Content type
All
85,188Total entries
1Added by human
85,187Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,805 results
Model Releases

SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes

DGX agent

arXiv:2606.01912v1 Announce Type: new Abstract: Smart homes are evolving toward complex state-dependent living environments, requiring Large Language Models (LLMs) to reason over user intent, preferen

model-releasesarxiv-cs-ai
2 Jun 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence

DGX agent

arXiv:2606.02380v1 Announce Type: cross Abstract: As LLM-based agents expand their operational scope, reliability becomes a prerequisite for real-world deployment. However, in practical applications,

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Sparse FEONet: A Low-Cost, Memory-Efficient Operator Network via Finite-Element Local Sparsity for Parametric PDEs

DGX agent

arXiv:2601.00672v2 Announce Type: replace-cross Abstract: In this paper, we study the finite element operator network (FEONet), an operator-learning method for parametric problems, originally introduc

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Spatially Distributed Task-Oriented Compression for Multi-Emitter Localization and Characterization with Spectral Overlap

DGX agent

arXiv:2606.01446v1 Announce Type: cross Abstract: Radio frequency spectrum awareness requires the ability to detect, localize, and characterize emitters in dense and contested wireless environments. I

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Spectra-Guided Neural Tucker Factorization

DGX agent

arXiv:2606.00584v1 Announce Type: cross Abstract: This paper proposes Spectra-Guided Neural Tucker Factorization (SG-NTF) for High-Dimensional and Incomplete (HDI) tensor completion. Circumventing dis

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Stability Analysis of Sharpness-Aware Minimization

DGX agent

arXiv:2301.06308v2 Announce Type: replace-cross Abstract: Sharpness-aware minimization (SAM) is a training method that seeks to find flat minima in deep learning, resulting in state-of-the-art perform

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Stable Velocity: A Variance Perspective on Flow Matching

DGX agent

arXiv:2602.05435v2 Announce Type: replace Abstract: While flow matching is elegant, its reliance on single-sample conditional velocities leads to high-variance training targets that destabilize optimi

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning

DGX agent

arXiv:2606.00148v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) often know the rule but pick the wrong answer: on abstract visual reasoning (AVR) tasks, a model can describe

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Strategizing at Speed: A Learned Model Predictive Game for Multi-Agent Drone Racing

DGX agent

arXiv:2602.06925v2 Announce Type: replace Abstract: Autonomous drone racing pushes the boundaries of high-speed motion planning and multi-agent strategic decision-making. Success in this domain requir

model-releasesarxiv-cs-ro
2 Jun 2026
Model Releases

StreamingVLM: Real-Time Understanding for Infinite Video Streams

DGX agent

arXiv:2510.09608v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) could power real-time assistants and autonomous agents, but they face a critical challenge: understanding near-i

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Structured Semantic Information Helps Retrieve Better Examples for In-Context Learning Applied to Few-Shot Relation Extraction

DGX agent

arXiv:2601.20803v2 Announce Type: replace Abstract: This paper presents several strategies to automatically obtain additional examples for in-context learning, effectively transforming relation extrac

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Subliminal Learning is a LoRA Artifact

DGX agent

arXiv:2606.00831v1 Announce Type: new Abstract: Subliminal learning is a phenomenon where language models can transmit behavioral traits to other models through seemingly innocuous data (Cloud et al.,

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in co…

DGX agent

Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier. First is our text foundation mode

model-releasesfireworks-ai--x
2 Jun 2026
Model Releases

SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory

DGX agent

arXiv:2606.00825v1 Announce Type: new Abstract: AI glasses present a compelling platform for AI agents to serve as personalized memory assistants. To be genuinely useful, such systems must move beyond

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Suppressing Forgery-Specific Shortcuts for Generalizable Deepfake Detection

DGX agent

arXiv:2606.01843v1 Announce Type: cross Abstract: Deepfake detection suffers from poor generalization across forgery methods, as existing models tend to rely on spurious method-specific shortcuts that

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Symbolic Neural Generation with Applications to Lead Discovery in Drug Design

DGX agent

arXiv:2510.23379v2 Announce Type: replace-cross Abstract: We investigate a relatively under-explored class of hybrid neurosymbolic models that integrate symbolic learning with neural reasoning to cons

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

T1: Tool-integrated Verification for Test-time Compute Scaling in Small Language Models

DGX agent

arXiv:2504.04718v2 Announce Type: replace-cross Abstract: Recent studies have demonstrated that test-time compute scaling effectively improves the performance of small language models (sLMs). However,

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

TabPrep: Closing the Feature Engineering Gap in Tabular Benchmarks

DGX agent

arXiv:2606.02384v1 Announce Type: new Abstract: Progress in tabular machine learning has largely focused on increasingly sophisticated model architectures. At the same time, feature engineering remain

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Taming System Complexity: Demystifying Software Engineering Agents in Diagnosing Linux Kernel Faults

DGX agent

arXiv:2505.19489v2 Announce Type: replace Abstract: The Linux kernel is a critical system, serving as the foundation for numerous systems. Bugs in the Linux kernel can cause serious consequences, affe

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Task diversity produces systematic transfer but inhibits continual reinforcement learning

DGX agent

arXiv:2606.00880v1 Announce Type: cross Abstract: Continual reinforcement learning aims to produce agents that learn not only to improve at their current tasks but also to adapt as task distributions

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

TCAR-Gen: Temporal Graph Retrieval with Evidence Fusion for Knowledge-Grounded Generation

DGX agent

arXiv:2606.00029v1 Announce Type: cross Abstract: Retrieval-augmented generation systems struggle with temporal reasoning and evidence fusion when answering complex questions over historical criminal

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

TECCI: Tricky Edits of Collected and Curated Images

DGX agent

arXiv:2606.01213v1 Announce Type: cross Abstract: Despite tremendous recent progress, current text-guided image editing methods still struggle with many aspects of editing involving instruction follow

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation

DGX agent

arXiv:2606.01031v1 Announce Type: cross Abstract: Audio-driven talking-head generation has advanced rapidly, yet existing evaluation protocols mainly rely on frame-wise metrics that assume strict temp

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

TempRet: Temporal Enhancement and Two-Stage Reranking for CVPR 2026 EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge

DGX agent

arXiv:2605.24470v2 Announce Type: replace Abstract: Video-text retrieval has witnessed remarkable progress driven by large-scale vision-language pretraining, yet most existing approaches inherit an im

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

TextFake: Benchmarking AI-Generated Image Detection on Text-Rich Images

DGX agent

arXiv:2606.01050v1 Announce Type: new Abstract: Recent AI-generated image (AIGI) detectors perform well on natural-image benchmarks, but their behavior on text-rich forgeries, such as fabricated scree

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

The Assistant as a Privileged Persona: A canonical reference in cross-persona self-recognition

DGX agent

arXiv:2606.00545v1 Announce Type: new Abstract: Post-trained language models can recognize their own outputs from a sentence or two out of context. In a companion paper itep{jack2026twomodes} we showe

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

The Case for Model Science: Verify, Explore, Steer, Refine

DGX agent

arXiv:2606.01189v1 Announce Type: new Abstract: We argue that the AI community is now ready to move beyond benchmarking and consolidate scattered efforts in model analysis into a systematic discipline

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing

DGX agent

arXiv:2606.02184v1 Announce Type: cross Abstract: These names do not exist. Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academ

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue

DGX agent

arXiv:2606.01901v1 Announce Type: cross Abstract: We introduce the Image Reconstruction Game, a fully automated benchmark in which a vision-language model issues corrective instructions to an image ge

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

The Invisible Coalition Partner: How LLMs Vote When Democracy Gets Concrete

DGX agent

arXiv:2606.00048v1 Announce Type: cross Abstract: Prior research has established that instruction-tuned large language models exhibit left-of-center political bias, measured exclusively through abstra

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models

DGX agent

arXiv:2605.05427v2 Announce Type: replace Abstract: Refusal rates are a poor proxy for LLM safety, i.e., a model may over-refuse benign prompts while still complying with harmful ones. We audit both f

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

The Shape of Wisdom: Decision Trajectories in Language Models

DGX agent

arXiv:2606.01202v1 Announce Type: new Abstract: Language models do not simply choose an answer at the output layer. In a 9,000-trajectory MMLU study across Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct,

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

This is actually one of the main advantages startups have over frontier labs, as long as there's a healthy spectrum of open-weight to closed…

DGX agent

This is actually one of the main advantages startups have over frontier labs, as long as there's a healthy spectrum of open-weight to closed-weight models on the cost-performance curve. Building a mod

model-releasesjerry-liu--x
2 Jun 2026
Model Releases

This is also available on the Claude Blog! https://claude.com/blog/a-harness-for-every-task-dynamic-workflows-in-claude-code

DGX agent

This post discusses dynamic workflows in Claude Code, enabling flexible task automation and execution patterns. The content is available on the official Claude Blog and addresses how Claude can be har

model-releasesthariq--x
2 Jun 2026
Model Releases

🌞This is big Local AI news! A new open-source Computer-Use LLM has just launched. Holo 3.1 is H Company’s (🇫🇷) new local computer-use age…

DGX agent

🌞This is big Local AI news! A new open-source Computer-Use LLM has just launched. Holo 3.1 is H Company’s (🇫🇷) new local computer-use agent model that beats Qwen3.5-397B, Kimi-K2.5, and Sonnet 4.6! Si

model-releasesclem-delangue--x
2 Jun 2026
Model Releases

this is fine 🐶☕️🔥

DGX agent

This post likely references the popular 'This is Fine' meme featuring a dog in a burning room, often used to comment on problematic situations being accepted or ignored. Without access to the specific

model-releasesemad-mostaque--x
2 Jun 2026
Model Releases

Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs

DGX agent

arXiv:2606.00592v1 Announce Type: new Abstract: Effective visual communication stems from the harmony of multiple design principles, such as readability, contrast, alignment, overlap, and coherence, w

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Time-Optimal Collision Avoidance Via a Greedy Polynomial Backward Sweep

DGX agent

arXiv:2606.01169v1 Announce Type: cross Abstract: Spacecraft collision avoidance for low-thrust satellites often requires determining not only how to maneuver, but also how late a maneuver can begin w

model-releasesarxiv-cs-ro
2 Jun 2026
Model Releases

TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning

DGX agent

arXiv:2606.01498v1 Announce Type: cross Abstract: Time series data inform critical decisions across many real-world domains. While large language model (LLM) agents can analyze data through natural la

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Tiny Recursive Models for Solving the J2-Perturbed Lambert Problem

DGX agent

arXiv:2606.00895v1 Announce Type: cross Abstract: This paper presents a fast, recursive neural solver for the J2-perturbed Lambert problem based on Tiny Recursive Models (TRM), termed the TRM-Perturbe

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

TLG: Temporal-Logic Grounding for Video Question Answering via Source-Annotation Reconstruction and Category-Targeted Reasoning

DGX agent

arXiv:2606.01591v1 Announce Type: new Abstract: The TimeLogic Challenge evaluates formal temporal-logic reasoning over video - 16 operators (before, after, until, since, always, co-occur, ordering, ..

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners

DGX agent

arXiv:2606.01810v1 Announce Type: new Abstract: Current benchmarks for embodied vision-language planning often favor linguistic next-token prediction over physically grounded next-state reasoning. Thi

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Topological Ignorability for Structural Causal Effects Beyond Means

DGX agent

arXiv:2606.01184v1 Announce Type: cross Abstract: Many interventions alter the structure of an outcome distribution rather than its mean: they can split a population into disconnected regimes, create

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Toward accurate RUL and SoH estimation using reinforced graph-based physics-informed neural networks enhanced with dynamic weights

DGX agent

arXiv:2507.09766v2 Announce Type: replace-cross Abstract: Accurate estimation of Remaining Useful Life (RUL) and State of Health (SoH) is essential for reliable Prognostics and Health Management (PHM)

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Towards Anytime Retrieval: A Benchmark for Anytime Person Re-Identification

DGX agent

arXiv:2509.16635v2 Announce Type: replace Abstract: In real applications, person re-identification (ReID) is expected to retrieve the target person at any time, including both daytime and nighttime, r

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Towards Lightweight Reliability: Using Soft Prompts for Hallucination Mitigation in Large Language Models

DGX agent

arXiv:2606.00919v1 Announce Type: new Abstract: Large language models (LLMs) have seen widespread adoption across various domains, yet their reliability is frequently undermined by hallucinations - re

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Towards Multidisciplinary Summarization of Hospital Stays: Efficient Sentence-Level Clinical Provenance Categorization

DGX agent

arXiv:2606.02487v1 Announce Type: new Abstract: Effective 'all-team' summarization in high-complexity settings like the Neonatal Intensive Care Unit (NICU) requires aggregating insights from diverse d

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Towards Simple and Provable Parameter-Free Adaptive Gradient Methods

DGX agent

arXiv:2412.19444v2 Announce Type: replace Abstract: Optimization algorithms such as AdaGrad and Adam have significantly advanced the training of deep models by dynamically adjusting the learning rate

model-releasesarxiv-cs-lg
2 Jun 2026
← Previous
1…236237238239240…476
Next →