AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,115
  • Agents7,313
  • Applications5,228
  • Concepts5
  • Hardware1,762
  • Industry6,105
  • Local Ai4,756
  • Model Releases22,759
  • Research19,333
  • Safety12,889
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,115
  • Agents7,313
  • Applications5,228
  • Concepts5
  • Hardware1,762
  • Industry6,105
  • Local Ai4,756
  • Model Releases22,759
  • Research19,333
  • Safety12,889
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
85,115Total entries
1Added by human
85,114Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,767 results
Model Releases

NebulaExp-8B: An Empirical Post-Training Pipeline via Full-Scale Ablation Research

DGX agent

arXiv:2606.26671v1 Announce Type: new Abstract: Post-training alignment determines the reasoning and human preference following capabilities of large language models, yet most existing works withhold

model-releasesarxiv-cs-ai
26 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Necessary but Not Sufficient: Temperature Control and Reproducibility in LLM-as-Judge Safety Evaluations

DGX agent

arXiv:2606.26185v1 Announce Type: new Abstract: LLM-as-judge ('grader') components are now standard in evaluation harnesses, including safety evaluations where a pass/fail verdict may gate downstream

model-releasesarxiv-cs-lg
26 Jun 2026
Model Releases

Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context

DGX agent

arXiv:2606.26493v1 Announce Type: new Abstract: Diffusion language models offer a promising alternative to autoregressive models due to their potential for parallel and iterative generation. However,

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

New paper on giving LLM agents experience that improves the weights and stays readable at the same time. Agent-experience methods split into…

DGX agent

New paper on giving LLM agents experience that improves the weights and stays readable at the same time. Agent-experience methods split into two camps. Externalized natural-language rules stay interpr

model-releasesdair-ai--x
26 Jun 2026
Model Releases

NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models

DGX agent

arXiv:2606.27047v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly technical dom

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

OI-Bench: An Option Injection Benchmark for Evaluating LLM Susceptibility to Directive Interference

DGX agent

arXiv:2601.13300v2 Announce Type: replace Abstract: Benchmarking large language models (LLMs) is critical for understanding their capabilities, limitations, and robustness. In addition to interface ar

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

OpenAI hopes to make GPT-5.6 generally available in the coming weeks and says 'this kind of government access process' should not become the long-term default (Amrith Ramkumar/Wall Street Journal)

DGX agent

Amrith Ramkumar / Wall Street Journal: OpenAI hopes to make GPT-5.6 generally available in the coming weeks and says “this kind of government access process” should not become the long-term default —

model-releasestechmeme
26 Jun 2026
Model Releases

OpenAI introduces GPT-5.6 to challenge Claude Mythos 5

DGX agent

OpenAI Group PBC today introduced GPT-5.6, a new series of large language models that it says can outperform Claude Mythos 5 across certain coding tasks. The most advanced algorithm in the lineup is k

model-releasessiliconangle
26 Jun 2026
Model Releases

OpenAI releases three versions of GPT-5.6, called Sol, Terra, and Luna, as a limited preview to ~20 companies, with participants disclosed to the US government (Axios)

DGX agent

Axios: OpenAI releases three versions of GPT-5.6, called Sol, Terra, and Luna, as a limited preview to ~20 companies, with participants disclosed to the US government — OpenAI is rolling out GPT-5.6 F

model-releasestechmeme
26 Jun 2026
Model Releases

OpenAI says GPT-5.6 Sol and Terra were capable of identifying vulnerabilities but were unable to execute autonomous, end-to-end attacks against hardened targets (OpenAI)

DGX agent

OpenAI: OpenAI says GPT-5.6 Sol and Terra were capable of identifying vulnerabilities but were unable to execute autonomous, end-to-end attacks against hardened targets — GPT-5.6 is a new family of th

model-releasestechmeme
26 Jun 2026
Model Releases

OpenAI unveils GPT-5.6 amid US AI regulatory drama

DGX agent

Less than 24 hours after news broke that OpenAI would stagger its next model release at the request of the Trump administration, that model, GPT-5.6, is here. On Friday, the company unveiled the limit

model-releasesthe-verge-ai
26 Jun 2026
Model Releases

OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents

DGX agent

arXiv:2606.26350v1 Announce Type: new Abstract: Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented across isolated tas

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

OpenRCA 2.0: From Outcome Labels to Causal Process Supervision

DGX agent

arXiv:2606.27154v1 Announce Type: new Abstract: Root cause analysis (RCA) poses a holistic test of LLM agentic capabilities, such as long-context understanding, multi-step reasoning, and tool use. How

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

Over-parameterization and Adversarial Robustness in Neural Networks: An Overview and Empirical Analysis

DGX agent

arXiv:2406.10090v3 Announce Type: replace Abstract: Thanks to their extensive capacity, over-parameterized neural networks exhibit superior predictive capabilities and generalization. However, having

model-releasesarxiv-cs-lg
26 Jun 2026
Model Releases

Parametric Generalized Adaptive Moment Features (PG-AMF) for Bearing Fault Diagnosis and Machine Health Monitoring

DGX agent

arXiv:2606.26317v1 Announce Type: cross Abstract: Accurate fault diagnosis of rolling element bearings in rotating machinery is considered essential for ensuring industrial safety and enabling predict

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

Parametric Open Source Games

DGX agent

arXiv:2606.27068v1 Announce Type: cross Abstract: Open-source game theory studies agents whose behavior may depend on one another's decision procedures, but most existing models use discrete or symbol

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

Patent Representation Learning via Self-supervision

DGX agent

arXiv:2511.10657v2 Announce Type: replace-cross Abstract: We study self-supervised patent representation learning with contrastive objectives. A standard baseline constructs positives by encoding the

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection

DGX agent

arXiv:2606.26552v1 Announce Type: cross Abstract: The rapid advancement of generative models presents a significant challenge to existing deepfake detection methods, particularly given the widespread

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs

DGX agent

arXiv:2606.26666v1 Announce Type: new Abstract: Autoregressive large language model (LLM) serving is increasingly limited by key-value (KV) cache movement rather than dense matrix multiplication. Mode

model-releasesarxiv-cs-lg
26 Jun 2026
Model Releases

PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing

DGX agent

arXiv:2606.26551v1 Announce Type: new Abstract: While instruction-based image editing, enabled by multi-modal generative models, has advanced significantly, existing benchmarks lack a comprehensive ev

model-releasesarxiv-cs-cv
26 Jun 2026
Model Releases

PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for Photorealistic Portrait Generation

DGX agent

arXiv:2606.26930v1 Announce Type: new Abstract: Reinforcement Learning like Group Relative Policy Optimization (GRPO) has significantly advanced text-to-image post-training. However, current methods o

model-releasesarxiv-cs-cv
26 Jun 2026
Model Releases

Post-Training Recipe, More Than Model Family, Shapes Multi-Agent LLM Conversational Behavior

DGX agent

arXiv:2606.20632v2 Announce Type: replace-cross Abstract: Multi-LLM systems use multiple language models to deliberate, judge each other's outputs, or coordinate as agents. Their value depends on the

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

ProvenAI: Provenance-Native Traces of Evidence in Generated Answers

DGX agent

arXiv:2606.26449v1 Announce Type: cross Abstract: Retrieval-augmented systems routinely present citations alongside generated answers, yet a citation does not confirm that the corresponding source mea

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

Quoting OpenAI

DGX agent

We're beginning a limited preview of the GPT‑5.6 series: Sol, our flagship model; Terra, a balanced model for everyday work; and Luna, a fast and affordable model. Terra has competitive performance to

model-releasessimon-willison
26 Jun 2026
Model Releases

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

DGX agent

arXiv:2606.26907v1 Announce Type: new Abstract: While text-to-image (T2I) models have achieved remarkable progress, they struggle with real-world requests that are often underspecified, implicit, or d

model-releasesarxiv-cs-cv
26 Jun 2026
Model Releases

R2D-RL: A RoboCup 2D Soccer Environment for Multi-Agent Reinforcement Learning

DGX agent

arXiv:2606.18786v2 Announce Type: replace Abstract: Robot soccer is a challenging testbed for multi-agent reinforcement learning because it combines partial observability, cooperative and adversarial

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

Real-Time Safety Evaluation of Human Arm Operations Using a Wrist-Mounted IMU with PSM System

DGX agent

arXiv:2502.09241v2 Announce Type: replace Abstract: This paper presents a novel approach to real-time safety monitoring in human-robot collaborative manufacturing environments through a wrist-mounted

model-releasesarxiv-cs-ro
26 Jun 2026
Model Releases

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP

DGX agent

arXiv:2606.26794v1 Announce Type: cross Abstract: CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text ali

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

RedVox: Safety and Fairness Gaps in Speech Models Across Languages

DGX agent

arXiv:2606.26968v1 Announce Type: new Abstract: Speech-capable models are increasingly deployed in real-world applications across languages. Yet their safety and fairness beyond English settings and u

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

Refusal Lives Downstream of Persona in Chat Models

DGX agent

arXiv:2606.26161v1 Announce Type: new Abstract: Linear directions in activation space have been identified for both refusal and persona traits in instruction-tuned chat models, but the two have been s

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models

DGX agent

arXiv:2510.09976v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models such as OpenVLA, Octo, and pi_0 have shown strong generalization by leveraging large-scale demonstrations, yet t

model-releasesarxiv-cs-lg
26 Jun 2026
Model Releases

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs

DGX agent

arXiv:2606.27369v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically rely on ground-truth answers to assign rewards, limiting their applica

model-releasesarxiv-cs-lg
26 Jun 2026
Model Releases

ReportLogic: Evaluating Logical Quality in Deep Research Reports

DGX agent

arXiv:2602.18446v2 Announce Type: replace-cross Abstract: Users increasingly rely on Large Language Models (LLMs) for Deep Research, using them to synthesize diverse sources into structured reports th

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

Representation Costs in Data Science: Foundations and the Quasi-Banach Spaces of Deep Neural Networks

DGX agent

arXiv:2606.14954v3 Announce Type: replace-cross Abstract: We develop a general framework for analyzing representation costs of parametric data-fitting methods through their parameter-space regularizer

model-releasesarxiv-cs-lg
26 Jun 2026
Model Releases

Reproducibility Study of 'AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models'

DGX agent

arXiv:2606.26783v1 Announce Type: cross Abstract: Fang et al. (2025) introduced a null-space constrained projection, named AlphaEdit, for locate-then-edit knowledge editing methods, theoretically guar

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

Ribbon: Scalable Approximation and Robust Uncertainty Quantification

DGX agent

arXiv:2606.27269v1 Announce Type: cross Abstract: Reliably quantifying predictive uncertainty is difficult for complex, high-dimensional, or misspecified models. Both fully Bayesian and bootstrap resa

model-releasesarxiv-cs-lg
26 Jun 2026
Model Releases

Rolling Shutter Relative Pose Estimation Made Practical

DGX agent

arXiv:2606.26863v1 Announce Type: new Abstract: Rolling shutter (RS) cameras equip virtually all consumer devices, yet RS-aware relative pose estimation has remained impractical: the state-of-the-art

model-releasesarxiv-cs-cv
26 Jun 2026
Model Releases

RoPEMover: Depth-Aware Object Relocation via Positional Embeddings

DGX agent

arXiv:2606.27332v1 Announce Type: new Abstract: Moving an object in a single image requires geometry-consistent spatial rearrangement, including handling occlusions, revealing previously unseen region

model-releasesarxiv-cs-cv
26 Jun 2026
Model Releases

RSPC: A Benchmark for Modeling Stress and Psychiatric Conditions in Digitally Mediated Relationships using Psychiatrist Annotations

DGX agent

arXiv:2606.27247v1 Announce Type: new Abstract: In NLP, mental health conditions are often modeled as isolated phenomena, without interpersonal context. We use Reddit posts about long-distance relatio

model-releasesarxiv-cs-lg
26 Jun 2026
Model Releases

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

DGX agent

arXiv:2606.14397v2 Announce Type: replace Abstract: As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabi

model-releasesarxiv-cs-lg
26 Jun 2026
Model Releases

SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

DGX agent

arXiv:2606.26901v1 Announce Type: cross Abstract: Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diver

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

Scalable AI-assisted Workflow Management for Detector Design Optimization Using Distributed Computing

DGX agent

arXiv:2603.30014v2 Announce Type: replace-cross Abstract: The Production and Distributed Analysis (PanDA) system, originally developed for the ATLAS experiment at the CERN Large Hadron Collider (LHC),

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

Scaling Multi-Reference Image Generation with Dynamic Reward Optimization

DGX agent

arXiv:2606.26947v1 Announce Type: cross Abstract: While personalized image generation has achieved remarkable progress, multi-reference image generation (MRIG) remains a challenging task. Most existin

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology

DGX agent

arXiv:2606.26563v1 Announce Type: cross Abstract: Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metad

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

SciFig: Towards Automating Editable Figure Generation for Scientific Papers

DGX agent

arXiv:2601.04390v2 Announce Type: replace Abstract: High-quality methodology figures are central to scientific communication, yet they remain difficult and time-consuming to create. Such figures must

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

Securing agentic AI with perimeter guardrails: What's new in VPC Service Controls

DGX agent

As enterprises scale autonomous AI agents into production, enabling safe innovation requires robust architectural guardrails. AI agents connect across tools and datasets, so it’s essential to establis

model-releasesgoogle-cloud-ai
26 Jun 2026
Model Releases

See & Sniff: Learning Visuo-Olfactory Representations

DGX agent

arXiv:2606.27307v1 Announce Type: new Abstract: While modern multimodal models integrate vision with language, audio, or touch, olfaction remains largely unexplored due to the lack of paired visuo-olf

model-releasesarxiv-cs-cv
26 Jun 2026
Model Releases

ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP

DGX agent

arXiv:2606.27027v1 Announce Type: cross Abstract: With the rapid evolution of LLM-driven agents, Model Context Protocol (MCP), an open protocol bridging LLMs with external tools, has quickly become fo

model-releasesarxiv-cs-ai
26 Jun 2026
← Previous
1…166167168169170…475
Next →