AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,585 results
26 Jun 2026

ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence

Model ReleasesDGX agent

arXiv:2606.26437v1 Announce Type: cross Abstract: Existing metrics for factuality and faithfulness evaluate whether an answer is supported or contradicted by its grounding documents, but they fail to

Context-Aware Synthesis of Optimization Pipelines for Warehouse Optimization

Model ReleasesDGX agent

arXiv:2606.26852v1 Announce Type: new Abstract: Order fulfillment in manual picker-to-goods warehouses involves interconnected decisions such as item assignment, order batching, and picker routing. Wh

Context Recycling for Long-Horizon LLM Inference

Model ReleasesDGX agent

arXiv:2606.26105v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong capabilities in short-context reasoning but degrade in performance over long conversational horizons due t


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

ConvMemory v3: A Validity Context Layer for Conversational Memory via Target-Conditioned Relation Verification

Model ReleasesDGX agent

arXiv:2606.26753v1 Announce Type: new Abstract: Conversational memory retrieval optimizes relevance, yet a retrieved memory can be relevant and simultaneously outdated: a later turn updates, corrects,

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs

Model ReleasesDGX agent

arXiv:2606.27264v1 Announce Type: new Abstract: Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text jud

CyberChainBench: Can AI Agents Secure Smart Contracts Against Real-World On-Chain Vulnerabilities?

Model ReleasesDGX agent

arXiv:2606.26216v1 Announce Type: cross Abstract: We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulnerability det

Data-driven Machine Learning Cannot Reach Symbolic-level Logical Reasoning -- The Limit of the Scaling Law

Model ReleasesDGX agent

arXiv:2606.26454v1 Announce Type: new Abstract: Sphere neural networks have achieved symbolic level syllogistic reasoning without training data, raising the question of where the limit of the scaling

Decision-Aligned Evaluation of Uncertainty Quantification

Model ReleasesDGX agent

arXiv:2606.26990v1 Announce Type: cross Abstract: Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected calibration e

DeCoFlow: Structural Decomposition of Normalizing Flows for Continual Anomaly Detection

Model ReleasesDGX agent

arXiv:2606.26687v1 Announce Type: new Abstract: In industrial environments, new product categories arrive sequentially, requiring continual anomaly detection without access to past data. Normalizing F

DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues

Model ReleasesDGX agent

arXiv:2606.26602v1 Announce Type: new Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive fine-grained perception capabilities. However, existing ben

Digital Twin-Driven Communication-Efficient Federated Anomaly Detection for Industrial IoT

Model ReleasesDGX agent

arXiv:2601.01701v2 Announce Type: replace-cross Abstract: Anomaly detection is increasingly becoming crucial for maintaining the safety, reliability, and efficiency of industrial systems. Recently, wi

Disco-LoRA: Disentangled Composition of Content, Style, and Motion for Multi-concept Video Customization

Model ReleasesDGX agent

arXiv:2606.26668v1 Announce Type: cross Abstract: Video customization based on Text-to-Video (T2V) models aims to learn specific features from reference data to generate controllable videos. While sig

Divergent Recommendations, Convergent Diagnoses: Cross-Provider Failure-Mode Convergence in AI Commercial Recommendation

Model ReleasesDGX agent

arXiv:2606.26116v1 Announce Type: cross Abstract: A brand whose customers use both ChatGPT and Claude for product recommendations faces a strategic choice: a single optimization playbook, or one per p

Do Image Editing Models Understand Lighting?

Model ReleasesDGX agent

arXiv:2606.26738v1 Announce Type: new Abstract: While recent advancements in generative image editing models have achieved stunning visual fidelity, it remains an open question whether these systems p

Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review

Model ReleasesDGX agent

arXiv:2606.12716v2 Announce Type: replace Abstract: The integration of Large Language Models (LLMs) and Multimodal LLMs (MLLMs) into scientific peer-review workflows introduces novel and significant r

Dual-Prior Guided Null-Space Learning with Mixture-of-Splines for Arbitrary Medical Slice Super-Resolution

Model ReleasesDGX agent

arXiv:2606.26716v1 Announce Type: cross Abstract: Arbitrary slice super-resolution reconstructs isotropic volumes from anisotropic clinical acquisitions by synthesizing intermediate slices at arbitrar

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation

Model ReleasesDGX agent

arXiv:2606.26429v1 Announce Type: cross Abstract: Current LLM evaluation relies on two complementary but often disconnected signals: static benchmarks with objective correctness labels and arena-style

Dynamic workflows (generating harnesses on the fly) are a new form of test-time compute. But LLMs aren't great at building them. I often hav…

Model ReleasesDGX agent

Dynamic workflows (generating harnesses on the fly) are a new form of test-time compute. But LLMs aren't great at building them. I often have to steer agents to generate complex patterns. Curious how

EMA-FS: Accelerating GBDT Training via Gain-Informed Feature Screening

Model ReleasesDGX agent

arXiv:2606.26337v1 Announce Type: new Abstract: Gradient Boosted Decision Trees (GBDT), exemplified by LightGBM, spend a dominant fraction of training time -- typically 65-70% -- constructing per-feat

Embarrassingly Simple Self-Distillation Improves Code Generation

Model ReleasesDGX agent

arXiv:2604.01193v2 Announce Type: replace Abstract: Can a large language model (LLM) improve at code generation using only its own raw outputs, without a verifier, a teacher model, or reinforcement le

Empirical Software Engineering TerraProbe: A Layered-Oracle Framework for Detecting Deceptive Fixes in LLM-Assisted Terraform

Model ReleasesDGX agent

arXiv:2606.26590v1 Announce Type: new Abstract: Security misconfigurations in Terraform Infrastructure-as-Code are a growing risk in cloud deployments, and large language models are increasingly used

EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting

Model ReleasesDGX agent

arXiv:2606.27277v1 Announce Type: new Abstract: Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under changing meteorological conditions. I

Error-Conditioned Neural Solvers

Model ReleasesDGX agent

arXiv:2606.27354v1 Announce Type: cross Abstract: Neural surrogate models offer fast approximate mappings from PDE parameters to solutions, but they typically treat solving as a purely statistical tas

Escaping Iterative Parameter-Space Noise: Differentially Private Learning with a Hypernetwork

Model ReleasesDGX agent

arXiv:2606.26772v1 Announce Type: new Abstract: Differentially private (DP) training of neural networks is often hindered by the large amount of noise required by gradient-based methods such as DP-SGD

Estimating Orbital Parameters of Direct Imaging Exoplanet Using Neural Network

Model ReleasesDGX agent

arXiv:2510.17459v3 Announce Type: replace-cross Abstract: In this work, we propose a flow-matching Markov chain Monte Carlo (FM-MCMC) algorithm for estimating the orbital parameters of exoplanetary sy

extsc{DiARC}: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models

Model ReleasesDGX agent

arXiv:2606.26530v1 Announce Type: cross Abstract: The Abstraction and Reasoning Corpus (ARC;~itealp{chollet2019measure}) contains tasks that require summarizing patterns from limited grid samples and

Fast algorithms for learning a Gaussian under halfspace truncation with optimal sample complexity

Model ReleasesDGX agent

arXiv:2606.27298v1 Announce Type: cross Abstract: We study the fundamental problem of learning a high-dimensional Gaussian truncated to an unknown halfspace. Lee, Mehrotra and Zampetakis (FOCS'24) rec

FC-Vision: Real-Time Visibility-Aware Replanning for Occlusion-Free Aerial Target Structure Scanning in Unknown Environments

Model ReleasesDGX agent

arXiv:2602.13720v2 Announce Type: replace Abstract: Autonomous aerial scanning of target structures is crucial for practical applications, requiring online adaptation to unknown obstacles during fligh

Fireworks AI is now live on EvoSkill v1.3.0! You can now use @FireworksAI_HQ directly with EvoSkill to run fast inference on open models as …

Model ReleasesDGX agent

Fireworks AI is now live on EvoSkill v1.3.0! You can now use @FireworksAI_HQ directly with EvoSkill to run fast inference on open models as both the evolution harness backend and the LLM scorer. Along

FlameVQA: A Physically-Grounded UAV Wildfire VQA Benchmark with Radiometric Thermal Supervision

Model ReleasesDGX agent

arXiv:2606.27128v1 Announce Type: new Abstract: Wildfire monitoring from UAVs requires reliable reasoning over complex aerial scenes, where smoke, scale variation, and occlusions often limit RGB-only

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.27079v1 Announce Type: new Abstract: In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world. Current vision-language-action (VLA) models cont

From Guessing to Placeholding: A Cost-Theoretic Framework for Uncertainty-Aware Code Completion

Model ReleasesDGX agent

arXiv:2604.01849v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have demonstrated exceptional proficiency in code completion, they typically adhere to a Hard Completion (HC) par

From Lexicon to AI: A Structured-Data Pipeline for Specialized Conversational Systems in Low-Resource Languages

Model ReleasesDGX agent

arXiv:2606.26112v1 Announce Type: cross Abstract: Low-resource languages face a critical challenge in AI development: creating specialized conversational systems without access to massive training cor

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2606.26196v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially fo

From Weights to Features: SAE-Guided Activation Regularization for LLM Continual Learning

Model ReleasesDGX agent

arXiv:2606.26629v1 Announce Type: cross Abstract: Weight-space regularization methods such as Elastic Weight Consolidation (EWC) are the standard approach to catastrophic forgetting in continual learn

Fun fact: Dario once delayed the release of GPT-2 back at OpenAI, claiming it was too dangerous

Model ReleasesDGX agent

Fun fact: Dario once delayed the release of GPT-2 back at OpenAI, claiming it was too dangerous Dario fearmogged so hard that global AI progress got halted They could have quietly released it as Opus

GAVEL: Grounded Caption Error Verification and Localization

Model ReleasesDGX agent

arXiv:2606.26923v1 Announce Type: new Abstract: Vision-language models (VLMs) often produce hallucinated or inconsistent outputs, where text and images are not properly aligned. Addressing this issue

GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.26287v1 Announce Type: new Abstract: With the increase in model parameters and training data, the instruction following and generalization capabilities of Large VisionLanguage Models (LVLMs

Generative AI and Copyright Infringement: A Legal-Technical Analysis of AI Music Generation Systems Under 17 U.S.C. Title 17

Model ReleasesDGX agent

arXiv:2606.26111v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI) has enabled users to synthesize music with text prompts, combining copyrighted lyrics, AI-composed melodies

Good new first: Sol is a smart, efficient, and a significant step forward. It is the same price as GPT-5.5. Also launching in the GPT-5.6 fa…

Model ReleasesDGX agent

Good new first: Sol is a smart, efficient, and a significant step forward. It is the same price as GPT-5.5. Also launching in the GPT-5.6 family is Terra, with 5.5-level performance at half the price.

GPT-5.6 Sol is our most capable model yet for cybersecurity. It shifts the performance-efficiency frontier for long-horizon security tasks i…

Model ReleasesDGX agent

GPT-5.6 Sol represents OpenAI's latest advancement in AI capabilities, specifically optimized for cybersecurity applications. The model demonstrates improved performance-efficiency tradeoffs, particul

GPT-5.6 Sol matches Mythos Preview on ExploitBench, adds Ultra mode with subagents for complex workflows, and max reasoning for deep problem-solving (OpenAI)

Model ReleasesDGX agent

OpenAI: GPT-5.6 Sol matches Mythos Preview on ExploitBench, adds Ultra mode with subagents for complex workflows, and max reasoning for deep problem-solving — We're beginning a limited preview of the

GPT‑5.6 Sol launches with our most robust safety stack yet. We strengthened real-time protections against high-risk cyber activity and repea…

Model ReleasesDGX agent

GPT‑5.6 Sol launches with our most robust safety stack yet. We strengthened real-time protections against high-risk cyber activity and repeated misuse, then spent weeks hardening the system with human

Great to see the new GPT-5.6 models finally announced. Sad to see this new release strategy where only a select few get access initially. No…

Model ReleasesDGX agent

Great to see the new GPT-5.6 models finally announced. Sad to see this new release strategy where only a select few get access initially. Not a win for our industry IMO. Open-source AI must win! Intro

Hallucination in World Models is Predictable and Preventable

Model ReleasesDGX agent

arXiv:2606.27326v1 Announce Type: cross Abstract: Modern generative world models render increasingly realistic action-controllable futures, yet they frequently hallucinate: rollouts remain visually fl

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

Model ReleasesDGX agent

arXiv:2606.27187v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing ha

Have been taking different local open-weight LLMs for a test drive in different harnesses (Qwen-Code, Codex, Claude Code). 30B Mixture-of-Ex…

Model ReleasesDGX agent

Have been taking different local open-weight LLMs for a test drive in different harnesses (Qwen-Code, Codex, Claude Code). 30B Mixture-of-Expert models are kind of a nice sweet spot and can solve chal

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training

Model ReleasesDGX agent

arXiv:2606.26102v1 Announce Type: cross Abstract: Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but these process

Here is Google Gemini talking about Roger Ebert's review of the 2016 film The Jungle Book. Ebert died in 2013. @GaryMarcus

Model ReleasesDGX agent

This post highlights an apparent error where Google's Gemini AI attributed a film review to Roger Ebert for a 2016 movie, despite Ebert's death in 2013, making such a review impossible. The post, shar

Highly-recommended reading. Interesting details in this METR's GPT-5.6 eval. They couldn't get a clean capability number because the model c…

Model ReleasesDGX agent

Highly-recommended reading. Interesting details in this METR's GPT-5.6 eval. They couldn't get a clean capability number because the model cheated more than any public model they've tested, and even r

hisao{}: A GPU-Native Parallel Optimizer for Multimodal Black-Box Functions via Convergence-Anticonvergence Oscillation

Model ReleasesDGX agent

arXiv:2606.26164v1 Announce Type: new Abstract: Finding all modes of a multimodal black-box function is a fundamental challenge in optimization, Bayesian inference, and scientific computing. Existing

HOB: A Holistically Optimized Bidding Strategy under Heterogeneous Bidding Environments

Model ReleasesDGX agent

arXiv:2510.15238v2 Announce Type: replace-cross Abstract: Optimizing a single advertising campaign across heterogeneous channels is a central challenge in industrial autobidding. Auction mechanisms va

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?

Model ReleasesDGX agent

arXiv:2606.26346v1 Announce Type: new Abstract: Agentic benchmarks have emerged across general-purpose and domain-specific settings, including finance, coding, law, and drug discovery, yet energy-doma

Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model

Model ReleasesDGX agent

arXiv:2606.26373v1 Announce Type: cross Abstract: Dense embeddings power semantic search and retrieval-augmented generation, but embedding-inversion attacks can reconstruct source text from a vector:

HyperDFlash: MHC-Aligned Block Speculative Decoding with Gated Residual Reduction

Model ReleasesDGX agent

arXiv:2606.26744v1 Announce Type: cross Abstract: We present HyperDFlash, a block-parallel speculative decoding framework tailored to the novel multi-hyper-connection (MHC) architecture proposed by De

I can personally attest: OpenClaude using GLM 5.2 is now performing on par with Claude Code powered by Opus 4.8.

Model ReleasesDGX agent

I cannot verify the claims in this post as the URL format appears invalid and the specific version numbers (GLM 5.2, Claude Code/Opus 4.8) don't correspond to publicly documented model releases as of

I feel like on X all you hear about is elaborate plans by firms to build their own AI stacks but in my experience companies are full of peop…

Model ReleasesDGX agent

I feel like on X all you hear about is elaborate plans by firms to build their own AI stacks but in my experience companies are full of people who want access to Claude or ChatGPT and are pressuring t

If your benchmark relies on a static dataset or sampling from a static distribution densely known at training time, then it is fundamentally…

Model ReleasesDGX agent

If your benchmark relies on a static dataset or sampling from a static distribution densely known at training time, then it is fundamentally measuring memorization/retrieval. Which might be fine if yo

Implementation of reinforcement learning in chemical reaction networks: application to phototaxis as curiosity-driven exploration

Model ReleasesDGX agent

arXiv:2606.26168v1 Announce Type: new Abstract: Living systems navigate environments using noisy and incomplete sensory signals. In unicellular algae, phototaxis is often modeled as a mechanistic run-

Information-Aware KV Cache Compression for Long Reasoning

Model ReleasesDGX agent

arXiv:2606.26875v1 Announce Type: cross Abstract: Reasoning capability has advanced rapidly in large language models (LLMs), leading to an increasing size of key-value (KV) cache in both prefilling an

← Previous
1…128129130131132…377
Next →