AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
All
85,136Total entries
1Added by human
85,135Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,778 results
Model Releases

The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?

DGX agent

arXiv:2606.04455v1 Announce Type: new Abstract: Current AI benchmarks evaluate agents on task execution within human-designed workflows. These evaluations fundamentally fail to measure a critical next

model-releasesarxiv-cs-ai
4 Jun 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

The new memory system will keep track of important details automatically. If you prefer the legacy saved memories experience, you can switch…

DGX agent

The new memory system will keep track of important details automatically. If you prefer the legacy saved memories experience, you can switch back in settings. The new memory system is rolling out to P

model-releasesopenai--x
4 Jun 2026
Model Releases

The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents

DGX agent

arXiv:2606.04296v1 Announce Type: new Abstract: As autonomous AI agents move from conversational systems to long-horizon software execution, runtime safety layers that decide when to interrupt an agen

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

The Variance Brain Foundation Models Forgot: Third-Order Statistics Predict Cognition Where Billion-Parameter Models Fail

DGX agent

arXiv:2606.04010v1 Announce Type: cross Abstract: Brain foundation models (BFMs) are self-supervised Transformers pretrained on fMRI data. We posit that these models should capture each subject's cogn

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Thinking Through Signs: PEEL as a Semiotic Scaffolding for Epistemically Accountable AI-Enabled Research

DGX agent

arXiv:2606.04152v1 Announce Type: new Abstract: Large language models are reshaping research practice while quietly eroding researchers epistemic accountability. This commentary introduces PEEL - Prot

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Today I'm launching a new project called SynthTraces 🔥 It is a minimal codebase to generate synthetic coding agent session traces using Pi …

DGX agent

Today I'm launching a new project called SynthTraces 🔥 It is a minimal codebase to generate synthetic coding agent session traces using Pi (from @badlogicgames) I wanted a large number of coding-agent

model-releasesclem-delangue--x
4 Jun 2026
Model Releases

Toward a Generalized Defense Across Sparse, Continuous, and Structured Parameter Attacks

DGX agent

arXiv:2606.04317v1 Announce Type: cross Abstract: Deep neural networks are increasingly deployed across heterogeneous and partially untrusted environments, where models are distributed through cloud s

model-releasesarxiv-cs-lg
4 Jun 2026
Model Releases

Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification

DGX agent

arXiv:2606.04037v1 Announce Type: new Abstract: Pre-deployment verification of enterprise artificial intelligence (AI) agents remains a critical gap between large language model (LLM) capability bench

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Toward Trustworthy Portrait Editing: Evaluation of Demographic Misrepresentation in I2I Models

DGX agent

arXiv:2602.16149v2 Announce Type: replace Abstract: Instruction-guided image-to-image (I2I) editors are increasingly used in consumer and professional visual workflows, where trustworthiness depends n

model-releasesarxiv-cs-cv
4 Jun 2026
Model Releases

Towards Efficient and Evidence-grounded Mobility Prediction with LLM-Driven Agent

DGX agent

arXiv:2606.05130v1 Announce Type: cross Abstract: Individual-level mobility prediction is central to urban simulation, transportation planning, and policy analysis. Supervised sequence models achieve

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

transitions like this are why we think it's helpful to have a provider-agnostic harness we used to talk more about swapping models when the …

DGX agent

transitions like this are why we think it's helpful to have a provider-agnostic harness we used to talk more about swapping models when the latest and greatest came out -- but the latest and greatest

model-releasesharrison-chase--x
4 Jun 2026
Model Releases

Treat Traffic Like Trees: A Semantic-Preserving Hierarchical Graph-Based Expert Framework for Encrypted Traffic Analysis

DGX agent

arXiv:2606.04517v1 Announce Type: cross Abstract: Graph-based deep learning methods have been widely employed in encrypted traffic analysis to exploit latent correlations across different granularitie

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Tree-Based Formalization of Multi-Agent Complementarity in Human-AI Interactions

DGX agent

arXiv:2606.04779v1 Announce Type: new Abstract: Complementarity is the case in which a human--AI interaction (HAI) outperforms the best prediction benchmark available among its members. Although this

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from k-Parity

DGX agent

arXiv:2601.22450v2 Announce Type: replace-cross Abstract: Masked Diffusion Language Models have recently emerged as a powerful generative paradigm, yet their generalization properties remain understud

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD

DGX agent

arXiv:2606.05058v1 Announce Type: cross Abstract: Computer-Aided Design (CAD) underpins modern engineering and manufacturing by enabling the creation of precise, editable 3D models. However, CAD resea

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Unifying Model-Free Efficiency and Model-Based Representations via Latent Dynamics

DGX agent

arXiv:2602.12643v2 Announce Type: replace-cross Abstract: We present Unified Latent Dynamics (ULD), a novel reinforcement learning algorithm that unifies the efficiency of model-free methods with the

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight LLMs

DGX agent

arXiv:2606.04035v1 Announce Type: cross Abstract: We present a systematic study of domain-dependent safety behavior in open-weight LLMs: 7 standardized experiments across 7 ethical domains, testing 5

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

VAMPS: Visual-Assisted Mathematical Problem Solving Benchmark

DGX agent

arXiv:2606.04244v1 Announce Type: new Abstract: Multimodal large language models are increasingly capable of complex reasoning, yet their performance often degrades when they must externalize a proble

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

VCIFBench: Evaluating Complex Instruction Following for Video Understanding

DGX agent

arXiv:2606.04588v1 Announce Type: new Abstract: Multimodal large language models have made rapid progress in video understanding, yet existing benchmarks largely rely on simple prompts and provide lim

model-releasesarxiv-cs-cl
4 Jun 2026
Model Releases

VGGSounder: Audio-Visual Evaluations for Foundation Models

DGX agent

arXiv:2508.08237v4 Announce Type: replace-cross Abstract: The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSound

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Video2LoRA: Parametric Video Internalization for Vision-Language Models

DGX agent

arXiv:2606.04351v1 Announce Type: cross Abstract: Processing video in vision-language models is expensive: each frame occupies hundreds of tokens, and inference cost scales with every frame and every

model-releasesarxiv-cs-cl
4 Jun 2026
Model Releases

We are excited to join Nvidia's Nemotron Coalition of leading AI labs working together to advance open frontier foundation models. To celebr…

DGX agent

We are excited to join Nvidia's Nemotron Coalition of leading AI labs working together to advance open frontier foundation models. To celebrate we have partnered with @nvidia and @nebiustf to provide

model-releasesnous-research--x
4 Jun 2026
Model Releases

We're building in Canada. 🇨🇦

DGX agent

We're building in Canada. 🇨🇦 For decades, Canada invested to build the research foundations that made modern AI possible. Now we have to build, train, and scale what comes next here at home. Canada's

model-releasescohere--x
4 Jun 2026
Model Releases

We're presenting ParseBench at CVPR 2026! ParseBench is the most comprehensive document understanding benchmark for VLMs. ✅ It contains 2k p…

DGX agent

We're presenting ParseBench at CVPR 2026! ParseBench is the most comprehensive document understanding benchmark for VLMs. ✅ It contains 2k pages of real-world enterprise documents ✅ It has comprehensi

model-releasesjerry-liu--x
4 Jun 2026
Model Releases

We're presenting ParseBench at CVPR 2026 today. 🦙 Come learn why document understanding is an AGI-complete problem (an agent can't act on a…

DGX agent

We're presenting ParseBench at CVPR 2026 today. 🦙 Come learn why document understanding is an AGI-complete problem (an agent can't act on a doc it can't correctly read, and reading a real enterprise t

model-releasesjerry-liu--x
4 Jun 2026
Model Releases

WETBench: A Benchmark for Detecting Task-Specific Machine-Generated Text on Wikipedia

DGX agent

arXiv:2507.03373v2 Announce Type: replace Abstract: Given Wikipedia's role as a trusted source of high-quality, reliable content, concerns are growing about the proliferation of low-quality machine-ge

model-releasesarxiv-cs-cl
4 Jun 2026
Model Releases

We’ve been researching new ways for ChatGPT memory to carry context across conversations and keep it useful over time. Today, that work is r…

DGX agent

We’ve been researching new ways for ChatGPT memory to carry context across conversations and keep it useful over time. Today, that work is rolling out as a more capable memory system in ChatGPT. https

model-releasesopenai--x
4 Jun 2026
Model Releases

What Are We Actually Benchmarking in Robot Manipulation?

DGX agent

arXiv:2606.04233v1 Announce Type: new Abstract: A robotics benchmark score measures success under one fixed evaluation setup, yet is routinely treated as evidence of general manipulation capability. W

model-releasesarxiv-cs-ro
4 Jun 2026
Model Releases

What happened when one of our models found a counterexample to an 80-year-old Erdős conjecture? Researchers @alexwei_, @HongxunWu, and @wjmz…

DGX agent

What happened when one of our models found a counterexample to an 80-year-old Erdős conjecture? Researchers @alexwei_, @HongxunWu, and @wjmzbmr1 shared the story on the OpenAI Podcast with @AndrewMayn

model-releasesopenai--x
4 Jun 2026
Model Releases

What If Prompt Injection Never Left? Exploring Cross-Session Stored Prompt Injection in Agentic Systems

DGX agent

arXiv:2606.04425v1 Announce Type: cross Abstract: Modern agentic systems transform LLMs from session-bounded assistants into stateful systems that persist and evolve shared world state across sessions

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

What's new for Managed Service for Apache Spark clusters

DGX agent

At Google Cloud, our goal is to let you run large-scale analytical and data science workloads with maximum efficiency so you can process big data pipelines, machine learning, and ETL tasks. We recentl

model-releasesgoogle-cloud-ai
4 Jun 2026
Model Releases

When Do Fewer Coordinates Suffice in DP-SGD?

DGX agent

arXiv:2606.04375v1 Announce Type: new Abstract: Differentially private stochastic gradient descent (DP-SGD) injects noise into every updated coordinate, making the injected noise energy scale with the

model-releasesarxiv-cs-lg
4 Jun 2026
Model Releases

When Seeing Is Not Believing -- A Benchmark for Search-Grounded Video Misinformation Detection

DGX agent

arXiv:2606.04098v1 Announce Type: new Abstract: Video misinformation increasingly operates at the semantic and evidential level: authentic footage may be selectively edited, temporally reordered, spli

model-releasesarxiv-cs-cv
4 Jun 2026
Model Releases

When you burn so much money you run out of options…

DGX agent

When you burn so much money you run out of options… Anthropic co-founder and President Daniela Amodei said the high cost of developing AI models is driving firms like hers to look to the public market

model-releasesgary-marcus--x
4 Jun 2026
Model Releases

With the new memory system, you can review and steer what ChatGPT remembers through a memory summary, with more visibility and control over …

DGX agent

OpenAI introduced a new memory system for ChatGPT that allows users to review and control what the AI remembers across conversations through a memory summary feature. This update provides enhanced tra

model-releasesopenai--x
4 Jun 2026
Model Releases

Wordle 1,810 4/6 ⬛🟨⬛⬛⬛ ⬛⬛🟨⬛🟨 🟨🟩⬛🟨🟨 🟩🟩🟩🟩🟩

DGX agent

This entry documents a Wordle game result where the player solved puzzle #1,810 in 4 attempts, using the color-coded feedback system (black for incorrect letters, yellow for correct letters in wrong p

model-releasesanthropic--x
4 Jun 2026
Model Releases

xAI has released a blog on Partnering with Vapi for Voice

DGX agent

xAI announced a partnership with Vapi to integrate voice capabilities into xAI's AI systems and services. The collaboration aims to enhance conversational AI by leveraging Vapi's voice technology plat

model-releaseselon-musk--x
4 Jun 2026
Model Releases

XSSR: Cross-Domain Self-Supervised Representative Selection for Efficient Annotation in Medical Image Segmentation

DGX agent

arXiv:2606.04301v1 Announce Type: new Abstract: Acquiring labeled medical image data is resource-intensive and a challenge further exacerbated in cross-domain scenarios where source and target dataset

model-releasesarxiv-cs-cv
4 Jun 2026
Model Releases

'Your AI Text is not Mine': Redefining and Evaluating AI-generated Text Detection under Realistic Assumptions

DGX agent

arXiv:2606.04906v1 Announce Type: cross Abstract: Although it is generally agreed that AI-generated text poses a broad societal risk, there is no common understanding in the AI-generated text detectio

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

20x Faster Training Data Reads with Alluxio and Ray Data: A Cross-Region Benchmark

DGX agent

This benchmark demonstrates how integrating Alluxio with Ray Data achieves 20x faster training data read speeds for cross-region machine learning workloads on Anyscale's platform. The study shows perf

model-releasesanyscale-ray
3 Jun 2026
Model Releases

95% token reduction. 30x faster execution. 90%+ task completion. Today at #MSBuild, we announced a major shift to move reasoning upstream: P…

DGX agent

95% token reduction. 30x faster execution. 90%+ task completion. Today at #MSBuild, we announced a major shift to move reasoning upstream: Pinecone Nexus now integrates directly with @Microsoft OneLak

model-releasespinecone--x
3 Jun 2026
Model Releases

A 3D Isovist World Model -- Revealing a City's Unseen Geometry and Its Emergent Cross-City Signature

DGX agent

arXiv:2606.03609v1 Announce Type: cross Abstract: Embodied agents that navigate cities rely on world models that predict how their surroundings will change as they move. But for navigation, what matte

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

A Benchmark for Semi-supervised Multi-modal Crowd Counting

DGX agent

arXiv:2606.03646v1 Announce Type: new Abstract: This paper constructs the first benchmark on semi-supervised multi-modal crowd counting. To lay the foundation for this unexplored task, we first formul

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

A Fast Methane Detection Pipeline on Board Satellites Based on Mag1c-SAS and LinkNet

DGX agent

arXiv:2606.03675v1 Announce Type: new Abstract: Methane is a potent greenhouse gas, and detecting leaks early via hyperspectral satellite imagery can help climate change mitigation efforts. Meanwhile,

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

A New Framework for Cybersecurity Refusals in AI Agents

DGX agent

arXiv:2606.02644v1 Announce Type: cross Abstract: Agentic scaffolds have dramatically improved LLM performance on complex, long-horizon tasks, yielding both broad benefits and amplified risks in domai

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

A Single-Loop Bilevel Deep Learning Method for Optimal Control of Obstacle Problems

DGX agent

arXiv:2601.04120v2 Announce Type: replace-cross Abstract: Optimal control of obstacle problems arises in a wide range of applications and is computationally challenging due to its nonsmoothness, nonli

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

A workflow audit is no longer the best way to figure out how to use AI in your job. Despite the advice from AI labs, I'm more convinced, bec…

DGX agent

A workflow audit is no longer the best way to figure out how to use AI in your job. Despite the advice from AI labs, I'm more convinced, because of AI's reasoning capabilities and long context horizon

model-releasesallie-k--miller--x
3 Jun 2026
Model Releases

Acceptance-Test-Driven Evaluation Protocols for Business-Centric LLM Systems

DGX agent

arXiv:2606.02755v1 Announce Type: cross Abstract: Large language model (LLM) applications are increasingly expected to satisfy deterministic institutional requirements while relying on probabilistic g

model-releasesarxiv-cs-ai
3 Jun 2026
← Previous
1…219220221222223…475
Next →