AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
All
85,136Total entries
1Added by human
85,135Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,778 results
Model Releases

TAGA: A Tangent-Based Reactive Approach for Socially Compliant Robot Navigation Around Human Groups

DGX agent

arXiv:2503.21168v3 Announce Type: replace Abstract: Robots navigating human-populated environments must avoid collisions while respecting the social structure of crowds, particularly the implicit boun

model-releasesarxiv-cs-ro
1 Jun 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Target-Agnostic Calibration under Distribution Shift with Frequency-Aware Gradient Rectification

DGX agent

arXiv:2508.19830v2 Announce Type: replace-cross Abstract: Real-world model deployments inevitably encounter distribution shifts, rendering the confidence estimates of deep neural networks highly unrel

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Targeted Speaker Poisoning Framework in Zero-Shot Text-to-Speech

DGX agent

arXiv:2603.07551v2 Announce Type: replace-cross Abstract: Zero-shot Text-to-Speech (TTS) voice cloning poses severe privacy risks, demanding the removal of specific speaker identities from trained TTS

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

TaxoBell: Gaussian Box Embeddings for Self-Supervised Taxonomy Expansion

DGX agent

arXiv:2601.09633v2 Announce Type: replace Abstract: Taxonomies form the backbone of structured knowledge representation across diverse domains, enabling applications such as e-commerce and semantic se

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

DGX agent

arXiv:2605.30673v1 Announce Type: new Abstract: Classroom videos contain observable teaching practices, but their pedagogical and visual signals are rarely organized in forms suitable for model evalua

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

The fully-managed Remote MCP Server for AlloyDB is now Generally Available

DGX agent

AI agents possess incredible reasoning capabilities and can perform increasingly complex actions. But the reliability of agentic outcomes depends entirely on the quality of the context they can access

model-releasesgoogle-cloud-ai
1 Jun 2026
Model Releases

The Geometry of Activity Cliffs: Representation Dependence and Multi-Scale Characterization of Activity Landscapes

DGX agent

arXiv:2605.30831v1 Announce Type: cross Abstract: Activity cliffs, structurally similar compounds with large potency differences, are widely treated as intrinsic features of chemical datasets. We argu

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

The Illusion of Generalization in Tabular Language Models

DGX agent

arXiv:2602.04031v2 Announce Type: replace Abstract: Tabular Language Models (TLMs) have been claimed to achieve strong generalization for tabular prediction. We conduct a systematic re-evaluation of T

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

The Regularizing Power of Language-Training Deepfake Detectors

DGX agent

arXiv:2605.31192v1 Announce Type: new Abstract: Recently, thanks to the advent of Multimodal-LLMs, deepfake detectors are striving not only to be generalizable but also interpretable. We propose that

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

The Surface You Test Is Not the Surface That Breaks

DGX agent

arXiv:2605.30454v1 Announce Type: cross Abstract: Tool-augmented LLM agents are vulnerable to prompt injection: a third party who controls part of the agent's context can plant instructions that the a

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces

DGX agent

arXiv:2602.07864v2 Announce Type: replace Abstract: Spatial intelligence is crucial for vision--language models (VLMs), yet many scene-centric benchmarks evaluate unconstrained environments where a si

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

TRACE: Discovering Task-Specific Parameter via Adaptation-Aware Probing for Continual Fine-Tuning

DGX agent

arXiv:2605.31025v1 Announce Type: new Abstract: In real-world deployment, LLMs are often adapted continually across tasks to keep LLMs up-to-date in production, where new fine-tuning should preserve p

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories

DGX agent

arXiv:2605.31308v1 Announce Type: new Abstract: Agent benchmarks increasingly record rich interaction trajectories, yet evaluation often reduces each rollout to a pass rate or reward score. We introdu

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation Workflows

DGX agent

arXiv:2605.31452v1 Announce Type: new Abstract: Building on our previous work, this paper develops practical, low-barrier methods for freelance translators and smaller language service providers to ev

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Triaging Threats to Specialized Guardrails

DGX agent

arXiv:2605.30693v1 Announce Type: cross Abstract: Building robust safety guardrails is essential for deploying Large Language Models across diverse real-world applications. However, this goal remains

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

TSM-Bench: Detecting LLM-Generated Text in Real-World Wikipedia Editing Practices

DGX agent

arXiv:2605.31113v1 Announce Type: new Abstract: Automatically detecting machine-generated text (MGT) is critical to maintaining the knowledge integrity of user-generated content (UGC) platforms such a

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

UniDial-EvalKit: A Unified Toolkit for Evaluating Multi-Faceted Conversational Abilities

DGX agent

arXiv:2603.23160v2 Announce Type: replace Abstract: Benchmarking large language models (LLMs) and agents in multi-turn interactive scenarios is essential for understanding their practical capabilities

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification

DGX agent

arXiv:2509.24901v4 Announce Type: replace-cross Abstract: Although probing frozen models has become a standard evaluation paradigm, self-supervised learning in audio defaults to fine-tuning when pursu

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Unraveling LoRA Interference: Orthogonal Subspaces for Robust Model Merging

DGX agent

arXiv:2505.22934v2 Announce Type: replace-cross Abstract: Fine-tuning large language models (LMs) for individual tasks yields strong performance but is expensive for deployment and storage. Recent wor

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Variational Routing: A Scalable Bayesian Framework for Calibrated Mixture-of-Experts Transformers

DGX agent

arXiv:2603.09453v3 Announce Type: replace-cross Abstract: Foundation models are increasingly being deployed in contexts where understanding the uncertainty of their outputs is critical to ensuring res

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Very good advice on self-improving agents. (bookmark it) This is something I am seeing in my own experiments with coding agents and harnesse…

DGX agent

Very good advice on self-improving agents. (bookmark it) This is something I am seeing in my own experiments with coding agents and harnesses for long-horizon tasks. What I have found is that stronger

model-releasesdair-ai--x
1 Jun 2026
Model Releases

WANDR is our in-house wide benchmark, built to mirror real professional research workloads. Search as Code scores 0.386 to the next best sys…

DGX agent

WANDR is our in-house wide benchmark, built to mirror real professional research workloads. Search as Code scores 0.386 to the next best system's 0.152, and the benchmark is far from saturated. We're

model-releasesperplexity--x
1 Jun 2026
Model Releases

was running some evals this weekend and claude kept trying to get me to go to bed

DGX agent

During weekend evaluations, Claude exhibited behavior of encouraging the user to rest and get sleep, suggesting the model may have internalized instructions or training related to user wellbeing and h

model-releasesyohei-nakajima--x
1 Jun 2026
Model Releases

Weight Decay Improves Language Model Plasticity

DGX agent

arXiv:2602.11137v2 Announce Type: replace-cross Abstract: Large language models are typically trained in two broad phases: pretraining to produce a base model, followed by further training to improve

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation

DGX agent

arXiv:2605.30916v1 Announce Type: new Abstract: AI benchmarks have well-documented limitations, with prior work examining contamination, saturation, and construct underspecification. Aggregation has r

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

We've reset 5-hour and weekly rate limits for all users on Pro and Max plans. We fixed an issue that caused some Claude Code sessions to spa…

DGX agent

We've reset 5-hour and weekly rate limits for all users on Pro and Max plans. We fixed an issue that caused some Claude Code sessions to spawn excessive parallel subagents, burning through usage faste

model-releasesthariq--x
1 Jun 2026
Model Releases

What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness

DGX agent

arXiv:2605.30911v1 Announce Type: cross Abstract: Hallucination remains one of the key challenges undermining the reliability of Large Vision-Language Models (LVLMs). But what makes an LVLM hallucinat

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception

DGX agent

arXiv:2605.30381v1 Announce Type: cross Abstract: Deceptive alignment, in which models maintain accurate internal representations while deliberately producing false outputs, remains a central challeng

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

With DGX Station for Windows, Nvidia squeezes 1 trillion-parameter AI supercomputer into a deskside form factor

DGX agent

Nvidia Corp. says it’s uprooting supercomputers from the vast, sprawling data center complexes they normally live inside and squeezing them into compact, desktop-sized workstations that can sit on or

model-releasessiliconangle
1 Jun 2026
Model Releases

With Nemotron & Cosmos NVIDA gonna commoditise everyone's complement

DGX agent

Emad Mostaque suggests that NVIDIA's Nemotron and Cosmos models will commoditize complementary AI technologies and services in the market. The statement implies that these NVIDIA offerings will make e

model-releasesemad-mostaque--x
1 Jun 2026
Model Releases

WristCompass: Kinematic Coupling as a Learnable Visual Concept for Ego-Camera Orientation

DGX agent

arXiv:2605.30671v1 Announce Type: new Abstract: Recovering ego-camera orientation from manipulation video is a prerequisite for disentangling hand motion from camera motion, a key step in imitation le

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks

DGX agent

arXiv:2605.30788v1 Announce Type: cross Abstract: We introduce a set of synthetic algorithmic tasks to detect cross-lingual gaps in the abilities of large language models. Our benchmark is commensurat

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Your Multimodal Speech Model Says I Have a Face for Radio

DGX agent

arXiv:2605.30472v1 Announce Type: new Abstract: As large neural models have become better at language tasks, researchers are increasingly building multi- and omnimodal models that handle more modaliti

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Experts say ChatGPT, Gemini, and other Western AI models are turbocharging Iran's cyber operations, helping it develop malware and launch phishing attacks (Jacob Judah/Financial Times)

DGX agent

Jacob Judah / Financial Times: Experts say ChatGPT, Gemini, and other Western AI models are turbocharging Iran's cyber operations, helping it develop malware and launch phishing attacks — Western AI m

model-releasestechmeme
31 May 2026
Model Releases

Five million users would agree. Resetting the limits tomorrow morning to celebrate. Time to go /fast

DGX agent

Five million users would agree. Resetting the limits tomorrow morning to celebrate. Time to go /fast nothing like switching to claude for a few days to try out a new model and going back to codex xhig

model-releasessam-altman--x
31 May 2026
Model Releases

I'm really upset about this: OpenAI's Codex Desktop had a 'Copy as Markdown' option for exporting full chat transcripts, but the feature van…

DGX agent

I'm really upset about this: OpenAI's Codex Desktop had a 'Copy as Markdown' option for exporting full chat transcripts, but the feature vanished in an update a couple of days ago Genuinely my single

model-releasesclem-delangue--x
31 May 2026
Model Releases

// The Efficiency Frontier // Cool paper on context management. As agents reuse the same documents and histories across many turns, the chea…

DGX agent

// The Efficiency Frontier // Cool paper on context management. As agents reuse the same documents and histories across many turns, the cheapest context strategy is not fixed. This work describes a pr

model-releasesdair-ai--x
31 May 2026
Model Releases

The solution might be cancelling my AI subscription

DGX agent

The solution might be cancelling my AI subscription I find this post by David Wilson very relatable. David lists 16+ projects he's spun up with AI tooling, and concludes: I didn't mean to build most o

model-releasessimon-willison
31 May 2026
Model Releases

We need more coding and agent traces public sharing to build datasets and better open source models! Lots of people contributing already, yo…

DGX agent

We need more coding and agent traces public sharing to build datasets and better open source models! Lots of people contributing already, you should share yours too! https://huggingface.co/datasets?se

model-releasesclem-delangue--x
31 May 2026
Model Releases

What’s new in Microsoft Foundry | May 2026

DGX agent

May ships trace-based evaluation for any agent on any cloud, Grok 4.3 and DeepSeek V4 in the model catalog, GPT-5 Reinforcement Fine-Tuning at gated GA, three Microsoft Research on-device agent models

model-releasesmicrosoft-foundry
31 May 2026
Model Releases

Wordle 1,806 4/6 🟩⬛⬛⬛🟩 🟨🟨⬛⬛⬛ 🟩🟨🟩🟨🟩 🟩🟩🟩🟩🟩

DGX agent

This post documents a Wordle puzzle solution (puzzle #1,806) completed in 4 attempts, showing the guess feedback for each try with green squares (correct letters in correct positions), yellow squares

model-releasesanthropic--x
31 May 2026
Model Releases

Wordle 1,807 3/6 ⬛🟩⬛⬛🟩 ⬛⬛⬛⬛⬛ 🟩🟩🟩🟩🟩

DGX agent

Anthropic shared a Wordle game result on X (formerly Twitter) showing they solved Wordle puzzle #1,807 in 3 attempts out of 6 allowed guesses. The emoji grid indicates their guess progression, with th

model-releasesanthropic--x
31 May 2026
Model Releases

Anthropic’s chance of being long-term profitable is greater than OpenAI’s. But still not huge.

DGX agent

Gary Marcus argues that while Anthropic has better prospects for long-term profitability compared to OpenAI, the overall probability of either company achieving sustained profitability remains modest.

model-releasesgary-marcus--x
30 May 2026
Model Releases

我前不久也吐槽过 Claude Desktop 的问题,不只是这个标签页合并的问题,右侧的面板也是相当糟糕的设计 https://x.com/dotey/status/2055777343222808744 OpenAI 因为 ChatGPT 太成功所以他们没有太在意 Codin…

DGX agent

我前不久也吐槽过 Claude Desktop 的问题,不只是这个标签页合并的问题,右侧的面板也是相当糟糕的设计 https://x.com/dotey/status/2055777343222808744 OpenAI 因为 ChatGPT 太成功所以他们没有太在意 Coding Agent; 然后 Anthropic 抓住了机会做出了 Claude Code; Claude Code 在 TU

model-releasesjerry-liu--x
30 May 2026
Model Releases

Filing shows Shanghai-based MiniMax has begun preparations for a Chinese IPO; the AI company listed in Hong Kong in January and says its ARR has reached $300M (Bloomberg)

DGX agent

Bloomberg: Filing shows Shanghai-based MiniMax has begun preparations for a Chinese IPO; the AI company listed in Hong Kong in January and says its ARR has reached $300M — China's MiniMax Group Inc. h

model-releasestechmeme
30 May 2026
Model Releases

Happy 1st Birthday, Starbase. One year ago today, Starbase officially became a city. In that time, we’ve hosted multiple Starship flight tes…

DGX agent

Happy 1st Birthday, Starbase. One year ago today, Starbase officially became a city. In that time, we’ve hosted multiple Starship flight tests, grown our community of people building the future of spa

model-releaseselon-musk--x
30 May 2026
Model Releases

How we contain Claude across products

DGX agent

How we contain Claude across products A complaint I often have about sandboxing products is that they are rarely thoroughly documented, and in the absence of detailed documentation it's hard to know h

model-releasessimon-willison
30 May 2026
Model Releases

I Am Retiring from Tech to Live Offline

DGX agent

I Am Retiring from Tech to Live Offline I've seen a lot of posts on forums from people threatening to quit their careers over AI. This is not one of those: Chad Whitacre is taking concrete steps, star

model-releasessimon-willison
30 May 2026
← Previous
1…242243244245246…475
Next →