AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “openai”

GridTimelineEvolution
2,581 results
2 Jun 2026

Microsoft’s first advanced reasoning AI is here

IndustryDGX agent

Microsoft announced a bunch of new in-house AI models at Build 2026, including a new 'flagship' model: MAI-Thinking-1. It's an ambitious step into model development for Microsoft, which introduced its

OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents

Model ReleasesDGX agent

arXiv:2606.02031v1 Announce Type: cross Abstract: Building capable visual web agents requires long-horizon reasoning, precise grounding, and robust interaction with dynamic real-world websites. Despit

Remind me, @Elonmusk, was GPT-5 really smarter than the smartest humans?

Model ReleasesDGX agent

Remind me, @Elonmusk, was GPT-5 really smarter than the smartest humans? @viktaur27 @Teslarati The rate of improvement from original GPT to GPT-3 is impressive. If this rate of improvement continues,

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Sensor Tower: ChatGPT has become the fastest app to hit 1B global MAUs by far; ChatGPT's MAUs are up 62% YoY in Q2 to date, Claude's MAUs are up 640% YoY to 56M (Harshita Mary Varghese/Reuters)

Model ReleasesDGX agent

Harshita Mary Varghese / Reuters: Sensor Tower: ChatGPT has become the fastest app to hit 1B global MAUs by far; ChatGPT's MAUs are up 62% YoY in Q2 to date, Claude's MAUs are up 640% YoY to 56M — Ope

1 Jun 2026

Evaluating using Mock Tool Calls to Quarantine Untrusted Prompt Inputs

ResearchDGX agent

arXiv:2605.30521v1 Announce Type: new Abstract: Large language models must frequently process untrusted inputs, such as judging an answer from another model or running tasks like spam and harm classif

Fingerprint launches AI Assistant Detection to spot traffic from ChatGPT, Gemini and Claude

Model ReleasesDGX agent

Device intelligence company FingerprintJS Inc. today launched a preview of two products built to identify traffic from artificial intelligence assistants, addressing a detection gap that has opened as

Incremental BPE Tokenization

ResearchDGX agent

arXiv:2605.30813v1 Announce Type: new Abstract: We propose a novel algorithm for incremental Byte Pair Encoding (BPE) tokenization. The algorithm processes each input byte in worst-case O(log^2 t) tim

One of the new, buzzy jobs in Silicon Valley is the AI Forward Deployed Engineer (FDE), an engineer who is embedded within a client organiza…

Model ReleasesDGX agent

One of the new, buzzy jobs in Silicon Valley is the AI Forward Deployed Engineer (FDE), an engineer who is embedded within a client organization to help customize solutions, such as building and tunin

31 May 2026

We need more coding and agent traces public sharing to build datasets and better open source models! Lots of people contributing already, yo…

Model ReleasesDGX agent

We need more coding and agent traces public sharing to build datasets and better open source models! Lots of people contributing already, you should share yours too! https://huggingface.co/datasets?se

weird the way this tweet was getting a ton of traffic and then just stopped. 🤷‍♂️

SafetyDGX agent

weird the way this tweet was getting a ton of traffic and then just stopped. 🤷‍♂️ I honestly think Elon’s best days are behind him: BYD is crushing Tesla in EVs. Waymo is crushing Tesla in AVs. Anthro

29 May 2026

BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base

Model ReleasesDGX agent

arXiv:2605.29379v1 Announce Type: new Abstract: We present BrahmicTokenizer-131K, a 131,072-vocabulary byte-level BPE tokenizer that closes the Brahmic compression gap at the 131K-vocabulary class whi

Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agents

AgentsDGX agent

arXiv:2605.29927v1 Announce Type: cross Abstract: Despite recent advances, LLM-based web agents still struggle with limited exploration, omission of critical steps, and sensitivity to task constraints

Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes

Model ReleasesDGX agent

arXiv:2605.28965v1 Announce Type: new Abstract: Linking free-text phenotype descriptions to ontology terms, typically referred to as phenotype annotation, is essential for the cross-study integration

LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis

Model ReleasesDGX agent

arXiv:2605.28876v1 Announce Type: cross Abstract: CI failure logs are large (median 5k lines, max 200k in this corpus) and noisy. Coding agents that try to debug them depend on an upstream tool to red

Today I learned that @Grimezsz has more courage in her pinky than Roon does in his entire cowardly body. Roon made excuses; Grimes defended …

SafetyDGX agent

Today I learned that @Grimezsz has more courage in her pinky than Roon does in his entire cowardly body. Roon made excuses; Grimes defended her views calmly and respectfully, like grownups should. Ope

28 May 2026

crazy that this was announced on the day tokenmaxxing died.

SafetyDGX agent

crazy that this was announced on the day tokenmaxxing died. Anthropic raised 65 billion in a funding round that valued the artificial intelligence company at 965 billion including the new investment,

Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems

SafetyDGX agent

arXiv:2605.27766v1 Announce Type: new Abstract: LLM safety evaluations predominantly test models in isolation, yet deployed AI agents increasingly operate within persistent social environments alongsi

Inversely Learning Transferable Rewards via Abstracted States

TutorialsDGX agent

arXiv:2501.01669v4 Announce Type: replace Abstract: Inverse reinforcement learning (IRL) has progressed significantly toward accurately learning the underlying rewards in both discrete and continuous

Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline

Model ReleasesDGX agent

arXiv:2605.27440v1 Announce Type: cross Abstract: Small changes to how a buyer phrases a question -- 'best CRM' vs 'top CRM' vs 'best CRM for a SaaS startup' -- produce substantially different brand r

Verifiable Benchmarking of Long-Horizon Spatial Biology

Model ReleasesDGX agent

arXiv:2605.28065v1 Announce Type: new Abstract: AI agents are increasingly useful for biological data analysis, but existing benchmarks mostly test broad biological knowledge, executable workflows, or

27 May 2026

AI training data provider Human Archive raises $8.2M

HardwareDGX agent

Artificial intelligence training data provider Human Archive Inc. today announced that it has raised 8.2 million in funding. Wing Venture Capital, NVP Capital, Y Combinator headlined the consortium th

E3: Issue-Level Backtesting for Automated Research Critique

Model ReleasesDGX agent

arXiv:2605.27072v1 Announce Type: cross Abstract: We present E3, an automated review assistant that augments reviewers and engineering teams by identifying decision-relevant technical concerns in rese

I really appreciate the lessons and technical ideas @samaysham & team were able to share about their tax agent system, which learns from pro…

AgentsDGX agent

I really appreciate the lessons and technical ideas @samaysham & team were able to share about their tax agent system, which learns from production traces to self-improve via detailed tracing tightly

“Tokens got burned for millions of dollars without any real significant ROI to show for it.” hearing this over and over again

SafetyDGX agent

“Tokens got burned for millions of dollars without any real significant ROI to show for it.” hearing this over and over again The same conversation is happening across tech right now and many of us sa

US law enforcement warns of 'anti-tech extremism' as AI hatred grows

IndustryDGX agent

U.S. law enforcement agencies are monitoring growing backlash to AI and have begun classifying anti-technology sentiment as an extremism threat, with unpublished reports from the Department of Homelan

26 May 2026

AI Content Moderation in Therapy Conversations

Model ReleasesDGX agent

arXiv:2605.25454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LL

Breaking the Chains of Probability: Neutrosophic Logic as a New Framework for Epistemic Uncertainty in Large Language Models

ResearchDGX agent

arXiv:2605.24053v1 Announce Type: new Abstract: Large Language Models (LLMs) are predominantly governed by probabilistic frameworks in which the sum of outcome probabilities is constrained to unity. T

D^2-Monitor: Dynamic Safety Monitoring for Diffusion LLMs via Hesitation-Aware Routing

Model ReleasesDGX agent

arXiv:2605.25893v1 Announce Type: new Abstract: Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monitoring

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection

SafetyDGX agent

arXiv:2509.13608v2 Announce Type: replace Abstract: As Large Multimodal Models (LMMs) become integral to daily digital life, understanding their safety architectures is a critical problem for AI Align

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Model ReleasesDGX agent

arXiv:2505.23764v3 Announce Type: replace-cross Abstract: Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, h

Nano Banana Pro vs. GPT Image 2 Nano Banana Pro wins on photorealism, 4K output, 14 reference image slots for product/scene consistency, and…

ApplicationsDGX agent

Nano Banana Pro vs. GPT Image 2 Nano Banana Pro wins on photorealism, 4K output, 14 reference image slots for product/scene consistency, and live web search for real-world accuracy. GPT Image 2 wins o

ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs

Model ReleasesDGX agent

arXiv:2507.10593v3 Announce Type: replace-cross Abstract: Every LLM tool call is structurally an RPC -- a function name, JSON arguments, and a serialized result -- yet each protocol (native Python, MC

25 May 2026

A look at the UK's AI Safety Institute, whose researchers probe AI models for safety gaps, as its work becomes a blueprint for other governments' AI policies (New York Times)

SafetyDGX agent

New York Times: A look at the UK's AI Safety Institute, whose researchers probe AI models for safety gaps, as its work becomes a blueprint for other governments' AI policies — The government's A.I. Se

An open source model has returned to #1 on the 3D Design leaderboard by Design Arena. Kimi K2.6 has reached the top of the leaderboard for 3…

Model ReleasesDGX agent

An open source model has returned to #1 on the 3D Design leaderboard by Design Arena. Kimi K2.6 has reached the top of the leaderboard for 3D Design, ahead of models 10X more expensive like Opus 4.7 b

Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering

Model ReleasesDGX agent

arXiv:2605.23497v1 Announce Type: new Abstract: Large language models are increasingly used for legal research, yet their fixed training cutoffs and reliance on static parametric knowledge are at odds

Reflex: Reinforcement Learning with Reflection Symmetry Exploitation in State-Based Continuous Control

SafetyDGX agent

arXiv:2605.23415v1 Announce Type: cross Abstract: Reinforcement learning has long struggled with poor sample efficiency. One promising approach to mitigate this problem is leveraging group-invariant M

23 May 2026

A Mechanistic Explanatory Strategy for XAI

Local AiDGX agent

arXiv:2411.01332v5 Announce Type: replace Abstract: Despite significant advancements in XAI, scholars note a persistent lack of solid conceptual foundations and integration with broader scientific dis

I once asked a bunch of e/accs how much damage was an acceptable risk relative to take (any) safety precautions. none answered. we may be ab…

SafetyDGX agent

I once asked a bunch of e/accs how much damage was an acceptable risk relative to take (any) safety precautions. none answered. we may be about to find out. One implication of the below is that we rea

it ain’t just me who sees the emperor has no clothes

SafetyDGX agent

it ain’t just me who sees the emperor has no clothes The AI bubble math doesn't add up. Anthropic spends 3 to make 1 and that’s before you include any and all other costs like staff or electricity. Mi

22 May 2026

Amplifying, Not Learning: Fine-Tuned AI Text Detectors Amplify a Pretrained Direction

SafetyDGX agent

arXiv:2605.21653v1 Announce Type: cross Abstract: AI text detectors amplify a pretrained typicality axis; they do not construct an AI-vs-human boundary. On raw encoders before any task supervision, pr

Code Researcher: Deep Research Agent for Large Systems Code and Commit History

Model ReleasesDGX agent

arXiv:2506.11060v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based coding agents have shown promising results on coding benchmarks, but their effectiveness on systems code rema

I truly miss the age of science and transparency in AI. Especially given how much money and political power and governance and scientific un…

Model ReleasesDGX agent

I truly miss the age of science and transparency in AI. Especially given how much money and political power and governance and scientific understanding is at stake. We don’t know for example • How man

21 May 2026

Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling

AgentsDGX agent

arXiv:2605.21470v1 Announce Type: new Abstract: Computer-use agents (CUA) automate tasks specified with natural language such as 'order the cheapest item from Taco Bell' by generating sequences of cal

can’t believe people assume that success on highly verifiable problems in math (where we don’t even know how many tests were performed and h…

SafetyDGX agent

can’t believe people assume that success on highly verifiable problems in math (where we don’t even know how many tests were performed and how many might have failed) iautomatically generalize to ever

proof too complicated, Claude help ELI5

Model ReleasesDGX agent

proof too complicated, Claude help ELI5 Today, we share a breakthrough on the planar unit distance problem, a famous open question first posed by Paul Erdős in 1946. For nearly 80 years, mathematician

🚀Qwen3.7-Max just landed at 56.6 on the Artificial Analysis Intelligence Index — a solid 4.8pt jump over Qwen3.6-Max-Preview. @ArtificialAn…

AgentsDGX agent

🚀Qwen3.7-Max just landed at 56.6 on the Artificial Analysis Intelligence Index — a solid 4.8pt jump over Qwen3.6-Max-Preview. @ArtificialAnlys ⚡️Sharper sci reasoning, stronger agentic chops, better c

20 May 2026

It’s a really special time to be alive…some thoughts from training this model 🧵

IndustryDGX agent

It’s a really special time to be alive…some thoughts from training this model 🧵 Today, we share a breakthrough on the planar unit distance problem, a famous open question first posed by Paul Erdős in

Once AI starts making solving open problems in novel ways it won’t stop. We are entering the final stage of human solutions to open problems…

IndustryDGX agent

Once AI starts making solving open problems in novel ways it won’t stop. We are entering the final stage of human solutions to open problems like this. Feels weird, doesn’t it? Today, we share a break

PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents

SafetyDGX agent

arXiv:2605.19932v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate over long and recurring external contexts, like document corpora and code repositories. Across in

the crazy part is that people are “clowning” me without knowing anything about the training or whether anything else other than scaled chang…

SafetyDGX agent

the crazy part is that people are “clowning” me without knowing anything about the training or whether anything else other than scaled changed or how the model does on anything else. (or what it costs

three of the things we are most excited about: 1. AGI accelerating research 2. AGI accelerating companies 3. personal AGI accelerating every…

IndustryDGX agent

three of the things we are most excited about: 1. AGI accelerating research 2. AGI accelerating companies 3. personal AGI accelerating everyone in achieving their goals today it was great to announce

19 May 2026

A Machine With Human-Like Memory Systems

Model ReleasesDGX agent

arXiv:2204.01611v3 Announce Type: replace Abstract: Inspired by the cognitive science theory, we explicitly model an agent with both semantic and episodic memory systems, and show that it is better th

Can LLMs Generate and Solve Linguistic Olympiad Puzzles?

Model ReleasesDGX agent

arXiv:2509.21820v2 Announce Type: replace Abstract: In this paper, we introduce a combination of novel and exciting tasks: the solution and generation of linguistic puzzles. We focus on puzzles used i

Causely: A Causal Intelligence Layer for Enterprise AI A Benchmark Study on SRE and Reliability Workflows

Model ReleasesDGX agent

arXiv:2605.18327v1 Announce Type: new Abstract: AI agents deployed into SRE workflows currently derive their understanding of environment state from raw observability telemetry at query time, paying a

Decart raises $300M for its AI optimization software, world models

HardwareDGX agent

Artificial intelligence developer Decart.ai Inc. today announced that it has raised 300 million in funding at a nearly 4 billion valuation. Radical Ventures led the round with participation from Nvidi

Episodic-Semantic Memory Architecture for Long-Horizon Scientific Agents

Model ReleasesDGX agent

arXiv:2605.17625v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into persistent scientific collaborators, context window saturation has emerged as a critical bottleneck. Scienti

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps

Model ReleasesDGX agent

arXiv:2605.17554v1 Announce Type: new Abstract: Frontier deep research agents (DRAs) plan a research task, synthesize across documents, and return a structured deliverable on demand. They are being de

EvilGenie: A Reward Hacking Benchmark

Model ReleasesDGX agent

arXiv:2511.21654v2 Announce Type: replace Abstract: We introduce EvilGenie, a benchmark for reward hacking in programming settings. We source problems from LiveCodeBench and create an environment in w

Fidelity Probes for Specification--Code Alignment

Model ReleasesDGX agent

arXiv:2605.17246v1 Announce Type: cross Abstract: We introduce fidelity probes: natural-language questions generated from a reference artifact with code-derived ground-truth answers, answered from a c

The last six months in LLMs in five minutes

Model ReleasesDGX agent

I put together these annotated slides from my five minute lightning talk at PyCon US 2026, using the latest iteration of my annotated presentation tool. # I presented this lightning talk at PyCon US 2

← Previous
1…3637383940…44
Next →