AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,553 results
15 Jul 2026

Repairing Shape-Prior Shortcuts in Long-Range Single-Shot Fringe Projection Profilometry

Model ReleasesDGX agent

arXiv:2607.11928v1 Announce Type: new Abstract: Single-shot fringe projection profilometry (FPP) networks that regress depth directly can exploit a shape-prior shortcut, recovering depth from object b

Reproducible Reservoir Computing with Thermally Driven Superparamagnets: Controlling Temperature Sensitivity

Model ReleasesDGX agent

arXiv:2607.12840v1 Announce Type: cross Abstract: Unconventional computing systems must demonstrate robust performance under real-world environmental conditions to enable practical deployments. We hav

Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2607.12985v1 Announce Type: new Abstract: Aligned language models routinely misreport under non-evidential incentive pressure: they agree with a confident user or overstate certainty even when t

Rethinking the Evaluation of Harness Evolution for Agents

Model ReleasesDGX agent

arXiv:2607.12227v1 Announce Type: new Abstract: We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness co

RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories

Model ReleasesDGX agent

arXiv:2512.04144v3 Announce Type: replace Abstract: Targeted interventions on language models, such as unlearning or model editing, aim to modify specific information, but their effects often propagat

RoboDesign1M: A Large-scale Dataset for Robot Design Understanding

Model ReleasesDGX agent

arXiv:2503.06796v2 Announce Type: replace Abstract: Robot design is a complex and time-consuming process that requires specialized expertise. Gaining a deeper understanding of robot design data can en

Sample Efficient Generative Optimization for Molecular Design

Model ReleasesDGX agent

arXiv:2607.12488v1 Announce Type: new Abstract: Molecular optimization in drug discovery, materials design, and catalysis requires searching vast chemical spaces under tight evaluation budgets, since

Scale-Aware Attention for Scarce Neural Data: An RG-Flow Transformer on Sleep-EDF EEG

Model ReleasesDGX agent

arXiv:2607.11950v1 Announce Type: cross Abstract: Brain field potentials are scale-free: their power spectra follow a 1/f^{eta} law whose aperiodic exponent eta tracks cortical state, and sleep depth

Scaling Point-in-Time Language Models

Model ReleasesDGX agent

arXiv:2607.11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromis

Self-Evolving In-Context Learning for Direct Pilot-to-Beamformer Design in MU-MISO Systems

Model ReleasesDGX agent

arXiv:2607.11970v1 Announce Type: cross Abstract: We develop an enhanced in-context learning (ICL) framework to improve the performance of pilot-based beamforming in multi-user multiple-input single-o

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

Model ReleasesDGX agent

arXiv:2607.12477v1 Announce Type: new Abstract: Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenar

Sensitivity to Subjective Expected Utility Maximization: A Methodological Study, with an Illustrative Application to LLM Decision-Making

Model ReleasesDGX agent

arXiv:2607.11920v1 Announce Type: cross Abstract: Evaluating decisions made under uncertainty is hard when labeled outcomes are scarce, costly, or confounded with luck. We treat subjective expected ut

SeqGPT: A Constrained Transformer Agent for the Inverse Designof Multi-Panel Composite Structures

Model ReleasesDGX agent

arXiv:2607.11910v1 Announce Type: cross Abstract: Optimizing composite stacking sequences to match continuous targets (e.g., Lamination or Buckling Parameters) with discrete manufacturing constraints

SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation

Model ReleasesDGX agent

arXiv:2506.12339v2 Announce Type: replace-cross Abstract: We present SheetMind, a modular multi-agent framework powered by large language models (LLMs) for spreadsheet automation via natural language

Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels

Model ReleasesDGX agent

arXiv:2607.12792v1 Announce Type: cross Abstract: Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach. Such evaluations, however, are se

SKooP: Symmetric Koopman Predictions for Faster and More Generalizable Legged Robot Locomotion with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.11624v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) algorithms classically suffer from poor sample efficiency. In robotics, a recent line of work has emerged addressi

So Many Opinions, So Many LLMs: Comparing Large Language Models to Traditional Machine Learning for Open- Ended Survey Analysis

Model ReleasesDGX agent

arXiv:2607.11890v1 Announce Type: cross Abstract: Open-ended surveys offer valuable insights, but they are notoriously difficult to analyze at scale. Building on previous work that employed traditiona

someone just told me about this take* on CUA this is one of those gell mann moments for me lol. i've been watching computer use since World …

Model ReleasesDGX agent

someone just told me about this take* on CUA this is one of those gell mann moments for me lol. i've been watching computer use since World of Bits (Shi, fan, karpathy, hernandez & liang 2017). we wer

Something I have been thinking about: in the past, the best engineers I knew spent a lot of time automating their work in various ways. Bett…

Model ReleasesDGX agent

Something I have been thinking about: in the past, the best engineers I knew spent a lot of time automating their work in various ways. Better vim/emacs automations, writing lint rules to catch repeat

Sophos launches Fusion, an AI-native ‘defense system’ to unify security tools

Model ReleasesDGX agent

Cybersecurity firm Sophos Ltd. today launched Sophos Fusion, a single platform that ties together its security operations, endpoint, network, identity, email and cloud protection. Sophos calls it the

Spherical-GOF: Geometry-Aware Panoramic Gaussian Opacity Fields for 3D Scene Reconstruction

Model ReleasesDGX agent

arXiv:2603.08503v2 Announce Type: replace Abstract: Omnidirectional images are increasingly used in robotics and vision due to their wide field of view. However, extending 3D Gaussian Splatting (3DGS)

SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise

Model ReleasesDGX agent

arXiv:2602.12783v3 Announce Type: replace-cross Abstract: Spoken query retrieval is an important interaction mode in modern information retrieval. However, existing evaluation datasets are often limit

tencent/Hy-Embodied-RxBrain-1.0 · Hugging Face

Model ReleasesDGX agent

Introduction RxBrain (Hy-Embodied-RxBrain-1.0) is a unified multimodal foundation model for embodied cognition — a single model that couples language reasoning with visual imagination to deliver three

TerraLogic: A Benchmark for Hierarchical Geospatial Reasoning in Earth Observation

Model ReleasesDGX agent

arXiv:2607.12497v1 Announce Type: new Abstract: Beyond perception, reasoning is essential in remote sensing for advanced interpretation, inference, and decision-making. Recent advances in large langua

TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

Model ReleasesDGX agent

arXiv:2607.13028v1 Announce Type: cross Abstract: Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground beh

The Benjamini--Hochberg Procedure Can Fail to Control the FDR for Correlated Two-Sided Gaussian Tests

Model ReleasesDGX agent

arXiv:2607.12208v1 Announce Type: cross Abstract: We show that the Benjamini--Hochberg procedure can fail to control the false discovery rate (FDR) at its nominal level for correlated two-sided Gaussi

The Blender MCP is now part of the Hermes Agent MCP Catalog! Easily activate and install the Blender MCP by running `hermes mcp install blen…

Model ReleasesDGX agent

The Blender MCP is now part of the Hermes Agent MCP Catalog! Easily activate and install the Blender MCP by running `hermes mcp install blender` and ask your agent to start using blender. Hermes will

The Capacity of Thought: Benchmarking Llama 3.2 in Semantic fMRI Neural Language Decoding and Improving the Huth Encoding-Model Baseline

Model ReleasesDGX agent

arXiv:2607.12079v1 Announce Type: new Abstract: Decoding continuous language from fMRI signals remains a core challenge in non-invasive brain-computer interface research. We present two complementary

The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context

Model ReleasesDGX agent

arXiv:2607.12963v1 Announce Type: new Abstract: As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by lo

The One-Word Census: Answer-Choice Conformity Across 44 Language Models

Model ReleasesDGX agent

arXiv:2607.12796v1 Announce Type: cross Abstract: When a language model must pick one answer from a large space of equally valid options, which does it pick -- and how often is it the same answer ever

The Risk of Exposed Cloud Functions and How to Harden

Model ReleasesDGX agent

Written by: Corné de Jong Introduction Mandiant security assessments frequently identify publicly exposed serverless applications that lack authentication, often as a result of specific business requi

The TopCoW Challenge -- Topology-Aware Circle of Willis Segmentation for CT and MR Angiography

Model ReleasesDGX agent

arXiv:2312.17670v5 Announce Type: replace Abstract: The Circle of Willis (CoW) is an important network of arteries connecting major circulations of the brain. Its vascular architecture is believed to

Thinking Machines just dropped a ~1T Omni model 🔥 > 1M context window, trained on 48T tokens of image, text, audio 🤯 > comes with a drafte…

Model ReleasesDGX agent

Thinking Machines just dropped a ~1T Omni model 🔥 > 1M context window, trained on 48T tokens of image, text, audio 🤯 > comes with a drafter, NVFP4 weights and Unsloth quants > transformers, llama.cpp

Thinky with a ~1T param, 41B active, apache-2 model Benchmarks are a clear step up from Nemotron Ultra (55B active), new best American model…

Model ReleasesDGX agent

Thinky with a ~1T param, 41B active, apache-2 model Benchmarks are a clear step up from Nemotron Ultra (55B active), new best American model, and omni input. A bit behind GLM 5.2 on agentic benchies,

This is for pretraining based on nemotron nvp4 figures, mfu etc Large context, multimodal, long RL we are seeing now could make it a multipl…

Model ReleasesDGX agent

This is for pretraining based on nemotron nvp4 figures, mfu etc Large context, multimodal, long RL we are seeing now could make it a multiple of this Worth noting how good the lite model is, similar t

this was such a fun talk with Cat and Simon, hope you get to check it out if you missed it

Model ReleasesDGX agent

this was such a fun talk with Cat and Simon, hope you get to check it out if you missed it 🆕This Year In Claude https://www.youtube.com/watch?v=uU5Gv2h8-9g @simonw chats with @_catwu and @trq212 about

🆕This Year In Claude https://www.youtube.com/watch?v=uU5Gv2h8-9g @simonw chats with @_catwu and @trq212 about the state of: - @claudeai Cod…

Model ReleasesDGX agent

🆕This Year In Claude https://www.youtube.com/watch?v=uU5Gv2h8-9g @simonw chats with @_catwu and @trq212 about the state of: - @claudeai Code - Claude Fable - @anthropicai culture & product strategy -

Token Reduction Is Not Cost Reduction

Model ReleasesDGX agent

arXiv:2607.12161v1 Announce Type: new Abstract: Context-reduction layers for API-based coding agents, including command-output compressors, retrieval rankers, and payload-optimizing proxies, are usual

Towards Self-Evolving Agents: A Human-Inspired Adaptive Exploration-Exploitation Framework for Genetic Network Programming

Model ReleasesDGX agent

arXiv:2607.11913v1 Announce Type: cross Abstract: Recent advancements in agentic AI have increasingly moved toward graph-based methods, driven by the demand for explainable, human-centered, and non-li

TRACE: An Operational Reasoning Schema for Auditable Agentic Commitments

Model ReleasesDGX agent

arXiv:2607.12480v1 Announce Type: new Abstract: This paper defines TRACE (Typed Reasoning And Commitment Evidence): a typed, versioned schema for recording reasoning traces, a reference procedure for

Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents

Model ReleasesDGX agent

arXiv:2607.12267v1 Announce Type: cross Abstract: Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy. We trace

Transforming LLMs into Efficient Cross-Encoders via Knowledge Distillation for RAG Reranking

Model ReleasesDGX agent

arXiv:2607.11933v1 Announce Type: new Abstract: Cross-encoders achieve high reranking accuracy in Retrieval-Augmented Generation (RAG) pipelines but impose quadratic inference costs that limit real-ti

Trending AND fastest-growing in the same month? Benchmarks change daily, but only @tryramp has the database of real receipts to track this. …

Model ReleasesDGX agent

Trending AND fastest-growing in the same month? Benchmarks change daily, but only @tryramp has the database of real receipts to track this. We can confirm: demand for open-weight inference and trainin

Tune in today at 10am PST!

Model ReleasesDGX agent

Tune in today at 10am PST! Livestream Alert: Run ComfyUI From Claude/Cursor with Comfy MCP Host: @PurzBeats Comfy MCP lets Claude, Cursor, Amp and almost any AI agent you're already using build, run,

UniVR: Thinking in Visual Space for Unified Visual Reasoning

Model ReleasesDGX agent

arXiv:2607.12800v1 Announce Type: new Abstract: Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation in

VanillaBench: The Hidden Accuracy Cost of Adversarial Robustness

Model ReleasesDGX agent

arXiv:2607.12545v1 Announce Type: cross Abstract: Adversarial robustness research has produced hundreds of defended models over the past decade, yet the literature almost universally reports robustnes

Verifier-Based Reinforcement Fine-Tuning of Reasoning Models for Thermal Energy Storage Control

Model ReleasesDGX agent

arXiv:2607.12856v1 Announce Type: new Abstract: Buildings are expected to shift cooling loads in response to grid conditions. Thermal energy storage (TES) enables this shift, but scheduling it well re

ViHoRec: A Quality-Controlled Vietnamese Hotel Recommendation Dataset and Cold-Start Benchmark

Model ReleasesDGX agent

arXiv:2607.12946v1 Announce Type: cross Abstract: Recommender-system research for Vietnamese remains limited by the absence of a public, well-documented hotel interaction resource. Building such a res

VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression

Model ReleasesDGX agent

arXiv:2607.12756v1 Announce Type: new Abstract: Vision-language models (VLMs) process large numbers of visual tokens, resulting in substantial inference latency and memory overhead. This has motivated

What Does a Temporal Benchmark Score Measure? Decomposing Channel Use in Video VLM Evaluation

Model ReleasesDGX agent

arXiv:2607.12304v1 Announce Type: new Abstract: A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions. 1.

When Close Enough Is Not Enough: Autoregressive Drift in Quantum Circuit Synthesis

Model ReleasesDGX agent

arXiv:2607.12780v1 Announce Type: cross Abstract: Quantum circuit optimization for fault-tolerant computing requires exact functional equivalence while minimizing expensive non-Clifford resources such

When Directional Accuracy Lies: A Base-Rate-Honest Benchmark for LoRA-Adapted TimesFM on Equity Forecasting

Model ReleasesDGX agent

arXiv:2607.12248v1 Announce Type: cross Abstract: Large pretrained time-series models such as TimesFM are attractive for financial forecasting, but raw directional accuracy is a misleading scoreboard

Who touches every token that flows through @anthropic? It’s not any one model, but it’s @katelyn_lesse, @angjiang and the platform team. A y…

Model ReleasesDGX agent

Who touches every token that flows through @anthropic? It’s not any one model, but it’s @katelyn_lesse, @angjiang and the platform team. A year ago, it was just a messages API. Today, their platform s

WikiSTAR: A System for Shedding Light on the Hidden History of Scientific Wikipedia Articles

Model ReleasesDGX agent

arXiv:2607.12441v1 Announce Type: new Abstract: Wikipedia plays a key role in shaping public understanding of science, and its openly accessible revision history is a unique record of how scientific k

X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

Model ReleasesDGX agent

arXiv:2607.12993v1 Announce Type: new Abstract: We present X-lens, a compact feed-forward model for metric depth estimation from a variable number of calibrated fisheye and pinhole views. To support r

xai-org/grok-build, now open source

Model ReleasesDGX agent

xai-org/grok-build, now open source xAI's grok CLI tool faced severe community backlash yesterday when it became apparent that running the command in a directory could upload that entire directory to

You don’t have to wait. Merch inspired by research & deployment. Available until sold out. https://openai.com/supply/

Model ReleasesDGX agent

OpenAI announced the launch of limited‑edition merchandise inspired by its research and deployment work, available for purchase until sold out via https://openai.com/supply/. The tweet highlighted “yo

Your coding agent doesn't need to leave the terminal to use Pinecone. We now ship official plugins and skills for the agentic IDEs and CLIs …

Model ReleasesDGX agent

Your coding agent doesn't need to leave the terminal to use Pinecone. We now ship official plugins and skills for the agentic IDEs and CLIs you are already building in: Claude Code, Cursor, GitHub Cop

14 Jul 2026

75.4% SWE Bench Verified / 53.9% SWE Bench Pro on 1 bit quantisation is 🤪 This is in line with my expectations & you can expect even lower …

Model ReleasesDGX agent

75.4% SWE Bench Verified / 53.9% SWE Bench Pro on 1 bit quantisation is 🤪 This is in line with my expectations & you can expect even lower drop off with NVP4 base trained models - why not run everythi

b10011

Model ReleasesDGX agent

server : refactor prompt cache state ownership (#25649) server : clear checkpoints upon prompt clear server : move the prompt state data to the server_prompt_cache Assisted-by: pi:llama.cpp/Qwen3.6-27

← Previous
1…8485868788…376
Next →