AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
85,136Total entries
1Added by human
85,135Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,778 results
Model Releases

Gram: Assessing sabotage propensities via automated alignment auditing

DGX agent

arXiv:2605.30322v1 Announce Type: cross Abstract: We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models ac

model-releasesarxiv-cs-ai
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents

DGX agent

arXiv:2605.29668v1 Announce Type: new Abstract: LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

GroundAct: Can LLM Agents Ground Actions in Environmental States?

DGX agent

arXiv:2508.05614v2 Announce Type: replace-cross Abstract: LLM agents achieve 85-96% success on tasks where instructions fully specify the action, but drop to 29-53% when action feasibility depends on

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human

DGX agent

arXiv:2605.28882v1 Announce Type: cross Abstract: With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important. However,

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

DGX agent

arXiv:2605.29218v1 Announce Type: new Abstract: Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing

DGX agent

arXiv:2605.29532v1 Announce Type: cross Abstract: Exploratory GUI testing is a particularly demanding setting for MLLM agents: without predefined test scripts, an agent must autonomously navigate an a

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Hallucination Detection-Guided Preference Optimization for Clinical Summarization

DGX agent

arXiv:2605.28910v1 Announce Type: cross Abstract: Large language models (LLMs) have shown promise on summarization tasks, but they often produce hallucinations, which are unsupported or incorrect stat

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching

DGX agent

arXiv:2605.29055v1 Announce Type: new Abstract: Hallucination remains a major reliability barrier for production LLM systems, particularly in multi-agent pipelines where unsupported claims can propaga

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Hands-on with Gemini Spark beta rolling out to AI Ultra subs: planned a birthday party from emails and calendar, but called a live-in boyfriend a 'close friend' (Reece Rogers/Wired)

DGX agent

Reece Rogers / Wired: Hands-on with Gemini Spark beta rolling out to AI Ultra subs: planned a birthday party from emails and calendar, but called a live-in boyfriend a “close friend” — Google's new AI

model-releasestechmeme
29 May 2026
Model Releases

Hear the architects of Gemini reflect on their journey to continue pushing the frontier of AI, on this episode of Release Notes. @JeffDean, …

DGX agent

Hear the architects of Gemini reflect on their journey to continue pushing the frontier of AI, on this episode of Release Notes. @JeffDean, @koraykv, @OriolVinyalsML, and @NoamShazeer sit down on came

model-releasesgoogle-ai--x
29 May 2026
Model Releases

HEART-Bench: Do LLM Agents Exhibit Human-like Psychology?

DGX agent

arXiv:2605.30058v1 Announce Type: new Abstract: While LLM agents have demonstrated remarkable task-oriented abilities such as planning, reasoning, and action, few works have treated them as complete h

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Here's an extended edit of the quote that includes a following fragment where Andrew Macdonald called the trade 'harder to justify' - full, …

DGX agent

Here's an extended edit of the quote that includes a following fragment where Andrew Macdonald called the trade 'harder to justify' - full, unedited transcript is here: https://gist.github.com/simonw/

model-releasessimon-willison--x
29 May 2026
Model Releases

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning

DGX agent

arXiv:2605.29782v1 Announce Type: cross Abstract: Reinforcement learning (RL) refines large language models (LLMs) by directly optimizing model behavior through reward signals. While accurate state va

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

How Braintrust turns customer requests into code with Codex

DGX agent

Braintrust leverages OpenAI's Codex model to automatically convert customer requests and natural language specifications into functional code, streamlining the software development process. This appli

model-releasesopenai
29 May 2026
Model Releases

How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions

DGX agent

arXiv:2605.29442v1 Announce Type: cross Abstract: AI coding agents increasingly act directly within software environments, yet existing analyses of their failures rely on benchmark trajectories that m

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

How Far Ahead Do LLMs Plan? Uncovering the Latent Horizon in Chain-of-Thought Reasoning

DGX agent

arXiv:2602.02103v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) reasoning has become a central mechanism for eliciting multi-step reasoning in Large Language Models (LLMs). Yet recent

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

How LoRA Remembers? A Parametric Memory Law for LLM Finetuning

DGX agent

arXiv:2605.30260v1 Announce Type: cross Abstract: Large Language Models (LLMs) must continuously learn and update knowledge to remain effective in dynamic real-world environments. While Low-Rank Adapt

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

How lucky are you to have been born when and where you are? Had Opus 4.8 in Claude Code whip up a new visualization of all humans who ever l…

DGX agent

How lucky are you to have been born when and where you are? Had Opus 4.8 in Claude Code whip up a new visualization of all humans who ever lived. In addition to being neat, it is an interesting test o

model-releasesethan-mollick--x
29 May 2026
Model Releases

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

DGX agent

arXiv:2605.30096v1 Announce Type: cross Abstract: Large language models (LLMs) can autonomously conduct multi-stage cyber attacks, but the consistency of their offensive behavior under repeated trials

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

I GOT THE DOMAIN! I FINALLY GOT IT!!!!!!!!!!1 🥳🎉 Paint​.NET is now at https://paint.net! Well, it will be just as soon as I push all the b…

DGX agent

I GOT THE DOMAIN! I FINALLY GOT IT!!!!!!!!!!1 🥳🎉 Paint​.NET is now at https://paint.net! Well, it will be just as soon as I push all the buttons to migrate content and set up redirects from getpaint​.

model-releasesanthropic--x
29 May 2026
Model Releases

- I still get nonstop questions about use cases, even though I think thinking about things in the use case framework is the wrong way of thi…

DGX agent

- I still get nonstop questions about use cases, even though I think thinking about things in the use case framework is the wrong way of thinking about it - Enterprises are starting to see actual busi

model-releasesallie-k--miller--x
29 May 2026
Model Releases

iLoRA: Bayesian Low-Rank Adaptation with Latent Interaction Graphs for Microbiome Diagnosis

DGX agent

arXiv:2605.30179v1 Announce Type: cross Abstract: Parameter-efficient adaptation has made LLMs practical for domain prediction, but standard LoRA still relies on a static low-rank update and does not

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

I’m going to let you in on a secret. I made a mistake once.* In August 2024.* It’s true. I actually got something wrong. I predicted that th…

DGX agent

I’m going to let you in on a secret. I made a mistake once.* In August 2024.* It’s true. I actually got something wrong. I predicted that there would be a collapse of the AI bubble (which in my judgem

model-releasesgary-marcus--x
29 May 2026
Model Releases

I'm suspicious of that that whole story about Uber blowing their AI budget and being disappointed in the results - I dug into it and it appe…

DGX agent

Simon Willison expresses skepticism about reports claiming Uber overspent on AI and was disappointed with the results, indicating he investigated the story and found issues with its accuracy or framin

model-releasessimon-willison--x
29 May 2026
Model Releases

Improving Adversarial Robustness of Attribution via Implicit Regularization

DGX agent

arXiv:2605.29983v1 Announce Type: cross Abstract: The adversarial robustness of attributions is a fundamental requirement for reliable explainability in deep learning, yet existing approaches typicall

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Improving Full Waveform Inversion in Large Model Era

DGX agent

arXiv:2603.00377v2 Announce Type: replace Abstract: Full Waveform Inversion (FWI) is a highly nonlinear and ill-posed problem that aims to recover subsurface velocity maps from surface-recorded seismi

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

In conversation with OpenAI’s @markchen90, Terence reflects on a future where AI reduces the cognitive friction of research, helps preserve …

DGX agent

In conversation with OpenAI’s @markchen90, Terence reflects on a future where AI reduces the cognitive friction of research, helps preserve the paths behind discovery, and expands what mathematicians

model-releasesopenai--x
29 May 2026
Model Releases

Indexing the Unreadable: LLM-Native Recursive Construction and Search of Service Taxonomies

DGX agent

arXiv:2605.29270v1 Announce Type: new Abstract: The era of the Internet of Agents (IoA) is taking shape: LLM agents are expected to fulfill user goals by orchestrating fast-growing populations of Mode

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Inferring the Size of Large Language Models From Popular Text Memorization

DGX agent

arXiv:2605.29223v1 Announce Type: new Abstract: The parameter counts of the most widely used large language models (LLMs) are often withheld by their developers, leaving model size -- a primary refere

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles

DGX agent

arXiv:2605.29473v1 Announce Type: cross Abstract: Language models are increasingly being deployed for conversational support in informal caregiving contexts, where interactions often extend beyond inf

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Data Agents

DGX agent

arXiv:2511.22884v2 Announce Type: replace Abstract: Data analysis has become an indispensable part of scientific research. To discover the latent knowledge and insights hidden within massive datasets,

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Interesting that the GPT-5 Pro series models have consistently been the best models for single-shot attempts at the hardest problems since l…

DGX agent

Interesting that the GPT-5 Pro series models have consistently been the best models for single-shot attempts at the hardest problems since last summer. There has been no real competition in all that t

model-releasesethan-mollick--x
29 May 2026
Model Releases

Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate

DGX agent

arXiv:2605.29889v1 Announce Type: cross Abstract: Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives

DGX agent

arXiv:2505.21627v4 Announce Type: replace-cross Abstract: State-of-the-art large language models require specialized hardware and substantial energy to operate. As a consequence, cloud-based services

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

I've largely switched over to using GPT-5.5 in recent weeks, which I like nearly as much as Opus 4.6 and 4.7, and is *very* reasonably price…

DGX agent

Jeremy Howard expresses positive views on GPT-5.5, stating he has recently switched to using it as his primary model and finds it nearly comparable to Anthropic's Opus 4.6 and 4.7 while offering signi

model-releasesjeremy-howard--x
29 May 2026
Model Releases

Just dropped 🧑‍🍳 Fixed version of DeepSeek-V4-Pro-NVFP4 by @NVIDIAAI https://huggingface.co/nvidia/DeepSeek-V4-Pro-NVFP4

DGX agent

NVIDIA has released a corrected version of DeepSeek-V4-Pro-NVFP4, a quantized model variant optimized for NVIDIA hardware using NV-FP4 (NVIDIA's 4-bit floating-point format). The model is available on

model-releasesclem-delangue--x
29 May 2026
Model Releases

K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance

DGX agent

arXiv:2605.29523v1 Announce Type: new Abstract: Large Language Models (LLMs) have advanced financial automation through Retrieval-Augmented Generation (RAG), yet hallucinations remain a critical barri

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing

DGX agent

arXiv:2605.29524v1 Announce Type: cross Abstract: Relay and reseller APIs increasingly intermediate access to large language models (LLMs), but users have no direct way to verify that a claimed endpoi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Kernel-based potential mean-field games with unbiased random Fourier U-statistics

DGX agent

arXiv:2605.29371v1 Announce Type: cross Abstract: We study the subclass of potential mean-field games in which the running interaction cost and the terminal target cost are both expressed through repr

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Kernel Renormalization in Bayesian Deep Neural Networks: the Equivalent Wishart Ansatz in the Proportional Regime

DGX agent

arXiv:2605.29684v1 Announce Type: new Abstract: The scaling limit where both the size of the training set P and the width N of a deep neural network grow at the same rate, the so-called proportional-w

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Knowledge Offloading: Decomposing LLMs into Sparse Backbones and Memory Modules

DGX agent

arXiv:2605.29075v1 Announce Type: new Abstract: LLMs encode both general capabilities and domain-specific knowledge in a single set of parameters. We ask whether this capacity can be reorganized: keep

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models

DGX agent

arXiv:2605.29459v1 Announce Type: new Abstract: Large language models route every input through a learned embedding table of shape |V| x d_model, consuming hundreds of millions to billions of trainabl

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Label-Free Reinforcement Learning via Cross-Model Entropy

DGX agent

arXiv:2605.29009v1 Announce Type: cross Abstract: Post-training large language models with reinforcement learning is bottlenecked by the reward signal. Existing approaches require either ground-truth

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Label Over Logic? How Source Cues Bias Human Fallacy Judgments More Than LLMs

DGX agent

arXiv:2605.29928v1 Announce Type: cross Abstract: As AI-generated and AI-assisted content floods online spaces, source labels attached to such content can distort human reasoning judgments, with downs

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Large language models reorganize representational geometry during in-context learning

DGX agent

arXiv:2605.28854v1 Announce Type: new Abstract: Large language models (LLMs) exhibit remarkable flexibility: they can adapt to novel tasks from in-context examples without any parameter updates, a cap

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Latent Performance Profiling of Large Language Models

DGX agent

arXiv:2605.30018v1 Announce Type: new Abstract: Large language models (LLMs) frequently achieve impressive scores on standardized benchmarks, yet accuracy alone offers a limited view of their capabili

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Leak@k: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding

DGX agent

arXiv:2511.04934v3 Announce Type: replace Abstract: Unlearning in large language models (LLMs) is critical for regulatory compliance and for building ethical generative AI systems that avoid producing

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Learning Design Skills as Memory Policies for Agentic Photonic Inverse Design

DGX agent

arXiv:2605.29421v1 Announce Type: new Abstract: Photonic crystal fiber (PCF) inverse design remains challenging because candidate geometries must satisfy coupled optical targets under expensive electr

model-releasesarxiv-cs-cl
29 May 2026
← Previous
1…246247248249250…475
Next →