AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “anthropic”

GridTimelineEvolution
1,724 results
5 Jun 2026

⚠️ The American public may be about to be fleeced. ⚠️ OpenAI, a deeply unethical company that was built on a series of lies, does not have a…

SafetyDGX agent

⚠️ The American public may be about to be fleeced. ⚠️ OpenAI, a deeply unethical company that was built on a series of lies, does not have a realistic road to profitability. And they reportedly want t

4 Jun 2026

OpenAI is in deep, deep trouble They are low on capital relative to the massive cash they are burning and there is only so much capital in t…

HardwareDGX agent

OpenAI is in deep, deep trouble They are low on capital relative to the massive cash they are burning and there is only so much capital in the world. No rational person would sell their Bitcoin or Nvi

Self-Reflective APIs: Structure Beats Verbosity for AI Agent Recovery


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
AgentsDGX agent

arXiv:2606.05037v1 Announce Type: cross Abstract: When an AI agent calls an API and hits a validation error, it needs more than what went wrong -- it needs what to do next. A self-reflective API retur

transitions like this are why we think it's helpful to have a provider-agnostic harness we used to talk more about swapping models when the …

Model ReleasesDGX agent

transitions like this are why we think it's helpful to have a provider-agnostic harness we used to talk more about swapping models when the latest and greatest came out -- but the latest and greatest

3 Jun 2026

ChatGPT makes history, becomes the fastest app to reach 1 billion monthly active users.

IndustryDGX agent

OpenAI's ChatGPT has crossed 1 billion global monthly active app users, becoming the fastest app ever to reach the milestone, according to estimates from Sensor Tower. The app reached this milestone i

If this prompt feels well written to you, it's because Suzanne is a writer in her little spare time! You can read her short story, Mall of A…

Model ReleasesDGX agent

If this prompt feels well written to you, it's because Suzanne is a writer in her little spare time! You can read her short story, Mall of America here: https://suzannewang.com/mall-of-america It's on

LAP: An Agent-to-Instrument Protocol for Autonomous Science

SafetyDGX agent

arXiv:2606.03755v1 Announce Type: new Abstract: Autonomous science is moving from demonstration to infrastructure. Large language model agents now plan experiments, and self-driving laboratories execu

Using open models and inference clouds (which serve open models) is a leading indicator of what is to come. The advantage of open weights is…

Model ReleasesDGX agent

Using open models and inference clouds (which serve open models) is a leading indicator of what is to come. The advantage of open weights is that you can train, serve, and continually improve your own

vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models

Model ReleasesDGX agent

arXiv:2603.04444v3 Announce Type: replace-cross Abstract: As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing -- se

2 Jun 2026

AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

Model ReleasesDGX agent

arXiv:2606.02240v1 Announce Type: cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, S

Connecting AI agents with unstructured data using Google Cloud Storage MCP Servers

Model ReleasesDGX agent

Google Cloud Storage (GCS) is a foundational component of the modern agentic tech stack and the preferred home for unstructured data at scale. As enterprises deploy agents in production, the critical

I wonder how many founders will pass on investors who passed on them in prior rounds I wonder how many would have three dinners & give them …

IndustryDGX agent

I wonder how many founders will pass on investors who passed on them in prior rounds I wonder how many would have three dinners & give them an allocation only to slash it to zero at the last moment. A

MazeBolt launches RADAR VectorAI to test enterprise defenses against AI-generated DDoS attacks

Model ReleasesDGX agent

Israeli distributed denial-of-service resilience company MazeBolt Technologies Ltd. today launched RADAR VectorAI, a new module that uses artificial intelligence to generate previously unseen DDoS att

ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL

SafetyDGX agent

arXiv:2606.01619v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) enables LLM agents to improve continuously from environment rewards, yet the resulting policies do not systematicall

1 Jun 2026

Join us, @mercor_ai, @Etched, and @AnthropicAI for a one-day hackathon in SF with a $50k top prize. Registrations close on 6/12. We can't wa…

AgentsDGX agent

Join us, @mercor_ai, @Etched, and @AnthropicAI for a one-day hackathon in SF with a 50k top prize. Registrations close on 6/12. We can't wait what to see you build! We're running a 24-hour hackathon J

One of the new, buzzy jobs in Silicon Valley is the AI Forward Deployed Engineer (FDE), an engineer who is embedded within a client organiza…

Model ReleasesDGX agent

One of the new, buzzy jobs in Silicon Valley is the AI Forward Deployed Engineer (FDE), an engineer who is embedded within a client organization to help customize solutions, such as building and tunin

Prediction: Nobody knows when this will all collapse, but 2026 will be remembered in hindsight as the year in which retail investors and ind…

SafetyDGX agent

Prediction: Nobody knows when this will all collapse, but 2026 will be remembered in hindsight as the year in which retail investors and index funds were left holding the bag. It's funny how people th

31 May 2026

weird the way this tweet was getting a ton of traffic and then just stopped. 🤷‍♂️

SafetyDGX agent

weird the way this tweet was getting a ton of traffic and then just stopped. 🤷‍♂️ I honestly think Elon’s best days are behind him: BYD is crushing Tesla in EVs. Waymo is crushing Tesla in AVs. Anthro

29 May 2026

even if @scaling01 turns out to be wrong about some of these, I respect the specificity.

SafetyDGX agent

even if @scaling01 turns out to be wrong about some of these, I respect the specificity. a bit more specific: - OpenAI will flourish -> meaning they will stay at the frontier and their market cap cont

Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes

Model ReleasesDGX agent

arXiv:2605.28965v1 Announce Type: new Abstract: Linking free-text phenotype descriptions to ontology terms, typically referred to as phenotype annotation, is essential for the cross-study integration

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

Model ReleasesDGX agent

arXiv:2605.30096v1 Announce Type: cross Abstract: Large language models (LLMs) can autonomously conduct multi-stage cyber attacks, but the consistency of their offensive behavior under repeated trials

i have strong reason to believe this is cope. we may find out soon…

SafetyDGX agent

i have strong reason to believe this is cope. we may find out soon… Im calling BS on this story. 1. That would be 100,000 employees spending 5k/mo each or 10,000 employees averaging 50k/mo each. No wa

This is a diary entry to myself, so I remember what AI was like today. It's just going to be a bullet-list stream of consciousness. - There …

Model ReleasesDGX agent

This is a diary entry to myself, so I remember what AI was like today. It's just going to be a bullet-list stream of consciousness. - There are still so many leaders that have never seen an agent run

this is on track to be this year’s “what did Ilya see?” rumor that people love that has nothing to do with reality.

SafetyDGX agent

this is on track to be this year’s “what did Ilya see?” rumor that people love that has nothing to do with reality. Im calling BS on this story. 1. That would be 100,000 employees spending 5k/mo each

28 May 2026

Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline

Model ReleasesDGX agent

arXiv:2605.27440v1 Announce Type: cross Abstract: Small changes to how a buyer phrases a question -- 'best CRM' vs 'top CRM' vs 'best CRM for a SaaS startup' -- produce substantially different brand r

Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit

Model ReleasesDGX agent

arXiv:2605.27439v1 Announce Type: cross Abstract: AI assistants like ChatGPT and Claude are recommendation engines, not search engines: they answer commercial queries by directly nominating brands rat

Stance Detection in Prediction Markets: Addressing Imbalanced Trader Commentary via Counterfactual Augmentation and Market Context

ResearchDGX agent

arXiv:2605.28745v1 Announce Type: new Abstract: Prediction markets such as Polymarket aggregate crowd beliefs into real-time probability estimates, and the comments traders post beneath each market co

Too many business leaders believe that AI says what it means. And it’s odd because we naturally attribute a high number of human traits to A…

Model ReleasesDGX agent

Too many business leaders believe that AI says what it means. And it’s odd because we naturally attribute a high number of human traits to AI, and yet we refuse to believe it can have hidden intent? 3

27 May 2026

E3: Issue-Level Backtesting for Automated Research Critique

Model ReleasesDGX agent

arXiv:2605.27072v1 Announce Type: cross Abstract: We present E3, an automated review assistant that augments reviewers and engineering teams by identifying decision-relevant technical concerns in rese

The Pope isn’t AGI-pilled

IndustryDGX agent

On Monday, Pope Leo XIV unveiled an encyclical letter addressing the societal implications of artificial intelligence. The letter, titled Magnifica Humanitas, warned that the 'use of AI is never a pur

“Tokens got burned for millions of dollars without any real significant ROI to show for it.” hearing this over and over again

SafetyDGX agent

“Tokens got burned for millions of dollars without any real significant ROI to show for it.” hearing this over and over again The same conversation is happening across tech right now and many of us sa

26 May 2026

BODHI: Precise OS Kernel Specification Inference

Model ReleasesDGX agent

arXiv:2605.23931v1 Announce Type: new Abstract: The formal verification of operating system kernels requires precise specifications that capture the intended behavior of system calls. Writing these sp

Demystifying the Mythos or Disrupting Bugonomics? From Zero-Day Asymmetry to Defender Remediation Throughput

TutorialsDGX agent

arXiv:2605.24632v1 Announce Type: cross Abstract: Recent demonstrations of large language models producing candidate and confirmed vulnerabilities in production software have renewed the narrative tha

Detectify debuts MCP server to let AI agents find and fix vulnerabilities in real time

AgentsDGX agent

Application security platform company Detectify AB today launched the Detectify MCP Server, a new integration layer that plugs the company’s security testing engines into artificial intelligence-drive

ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs

Model ReleasesDGX agent

arXiv:2507.10593v3 Announce Type: replace-cross Abstract: Every LLM tool call is structurally an RPC -- a function name, JSON arguments, and a serialized result -- yet each protocol (native Python, MC

> … [W]e keep finding things that are mysterious, even unsettling. We find structures that mirror results from human neuroscience. We find e…

ToolsDGX agent

> … [W]e keep finding things that are mysterious, even unsettling. We find structures that mirror results from human neuroscience. We find evidence of introspection. We find internal states that funct

25 May 2026

Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering

Model ReleasesDGX agent

arXiv:2605.23497v1 Announce Type: new Abstract: Large language models are increasingly used for legal research, yet their fixed training cutoffs and reliance on static parametric knowledge are at odds

Microsoft economist's hot take: Let it burn first

IndustryDGX agent

This post discusses how OpenAI is burning 17 billion annually despite 20+ billion in revenue because unit economics for AI inference are fundamentally broken, with the company losing more money on que

23 May 2026

100%: if OpenAI flops, the gravy train stops.

SafetyDGX agent

100%: if OpenAI flops, the gravy train stops. @GaryMarcus if OpenAI's IPO flops, gravy train is OVER! whole industry knows it, they're terrified of it, particularly as they're starting to get just how

A Mechanistic Explanatory Strategy for XAI

Local AiDGX agent

arXiv:2411.01332v5 Announce Type: replace Abstract: Despite significant advancements in XAI, scholars note a persistent lack of solid conceptual foundations and integration with broader scientific dis

I called both missteps and that’s part of why OpenAI and their shills hate me. 🤷‍♂️

Model ReleasesDGX agent

I called both missteps and that’s part of why OpenAI and their shills hate me. 🤷‍♂️ @GaryMarcus It all went wrong with GPT-5 and Sora 2 IMO. Until then OpenAI could do no wrong in the eyes of most peo

I once asked a bunch of e/accs how much damage was an acceptable risk relative to take (any) safety precautions. none answered. we may be ab…

SafetyDGX agent

I once asked a bunch of e/accs how much damage was an acceptable risk relative to take (any) safety precautions. none answered. we may be about to find out. One implication of the below is that we rea

Last I checked, the @llama_index OSS framework integrates with 78 (!!) vector stores. Investors (and builders starting out, and honestly me …

AgentsDGX agent

Last I checked, the @llama_index OSS framework integrates with 78 (!!) vector stores. Investors (and builders starting out, and honestly me for a while) deemed this space as commoditized. One of the l

my guess as to why:

SafetyDGX agent

my guess as to why: the OpenAI crowd is suddenly coming after me relentlessly, but too chicken to actual face me in a moderated debate. here’s why: a. the IPO is coming, but OpenAI’s has lost their le

22 May 2026

“Gary Marcus, former CEO of Geometric Intelligence and New York University professor, joins @SquawkStreet @CNBC to discuss why he still beli…

SafetyDGX agent

“Gary Marcus, former CEO of Geometric Intelligence and New York University professor, joins @SquawkStreet @CNBC to discuss why he still believes OpenAI could be the ‘WeWork’ of AI, why he finds Anthro

This is one of the things I dislike about managed agents. Is it the best DX? yes. Is it now much, much more usable because it's bring your o…

AgentsDGX agent

This is one of the things I dislike about managed agents. Is it the best DX? yes. Is it now much, much more usable because it's bring your own sandbox? yes (most startups now have Sandbox credits and

💯: Why OpenAI keeps taking childish shots at me, is exactly what @FrankRundatz says below: “you have the added audacity of being influentia…

SafetyDGX agent

💯: Why OpenAI keeps taking childish shots at me, is exactly what @FrankRundatz says below: “you have the added audacity of being influential enough to move the needle on the timing of OpenAI’s record-

21 May 2026

🚀Qwen3.7-Max just landed at 56.6 on the Artificial Analysis Intelligence Index — a solid 4.8pt jump over Qwen3.6-Max-Preview. @ArtificialAn…

AgentsDGX agent

🚀Qwen3.7-Max just landed at 56.6 on the Artificial Analysis Intelligence Index — a solid 4.8pt jump over Qwen3.6-Max-Preview. @ArtificialAnlys ⚡️Sharper sci reasoning, stronger agentic chops, better c

20 May 2026

Microsoft Senior AI developer just showed how they build AI agents with Claude at Microsoft. 34-minutes. free. By Microsoft team Opus 4.7 + …

Model ReleasesDGX agent

Microsoft Senior AI developer just showed how they build AI agents with Claude at Microsoft. 34-minutes. free. By Microsoft team Opus 4.7 + 1,400+ pre-built MCP tools plug Claude into agent → give it

The Growing Pains of Frontier Models: When Leaderboards Stop Separating and What to Measure Next

Model ReleasesDGX agent

arXiv:2605.18840v1 Announce Type: cross Abstract: Leaderboards rank frontier models on independent axes but do not reveal whether capabilities reinforce or trade off across releases -- and at the fron

19 May 2026

As companies approach AGI it would be illogical for most not to go and work there. AGI is a much bigger deal than most people still seem to …

ApplicationsDGX agent

As companies approach AGI it would be illogical for most not to go and work there. AGI is a much bigger deal than most people still seem to believe & it is only by most forecasts a few years away (if

Breaking: Andrej Karpathy just filed his Q1 2026 13F. Here's everything you need to know about his recent 13F Top 10 positions: 1. VanEck Se…

HardwareDGX agent

Breaking: Andrej Karpathy just filed his Q1 2026 13F. Here's everything you need to know about his recent 13F Top 10 positions: 1. VanEck Semiconductor ETF - SMH - [Put] — 2.04B 2. Nvidia - NVDA - [Pu

CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery

AgentsDGX agent

arXiv:2604.01658v2 Announce Type: replace Abstract: Large language model (LLM)-based evolution is a promising approach for open-ended discovery, where progress requires sustained search and knowledge

Episodic-Semantic Memory Architecture for Long-Horizon Scientific Agents

Model ReleasesDGX agent

arXiv:2605.17625v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into persistent scientific collaborators, context window saturation has emerged as a critical bottleneck. Scienti

EvilGenie: A Reward Hacking Benchmark

Model ReleasesDGX agent

arXiv:2511.21654v2 Announce Type: replace Abstract: We introduce EvilGenie, a benchmark for reward hacking in programming settings. We source problems from LiveCodeBench and create an environment in w

Fidelity Probes for Specification--Code Alignment

Model ReleasesDGX agent

arXiv:2605.17246v1 Announce Type: cross Abstract: We introduce fidelity probes: natural-language questions generated from a reference artifact with code-derived ground-truth answers, answered from a c

Implementing programmatic tool calling on Amazon Bedrock

IndustryDGX agent

In this post, we show three ways to implement Programmatic tool calling (PTC) on Amazon Bedrock: a self-hosted Docker sandbox on ECS for maximum control, a managed solution using Amazon Bedrock AgentC

The Range Shrinks, the Threat Remains: Re-evaluating LLM Package Hallucinations on the 2026 Frontier-Model Cohort

Model ReleasesDGX agent

arXiv:2605.17062v1 Announce Type: cross Abstract: Spracklen et al. (USENIX Security '25) showed that code-generating large language models hallucinate package names that do not exist on PyPI or npm at

[video] why we need a new continuity layer for long-running agents (claude did this video! all except the voice which was @elevenlabs)

Model ReleasesDGX agent

This video discusses the architectural need for a continuity layer in long-running AI agents, explaining how agents require persistent memory and state management mechanisms to maintain coherence acro

18 May 2026

Finally a semi-useful read on Mythos that is free of myth and talks about what this means more practically (not this is the end of the world…

TutorialsDGX agent

Finally a semi-useful read on Mythos that is free of myth and talks about what this means more practically (not this is the end of the world as we know it, but how do we deal with faster patches and a

← Previous
1…2324252627…29
Next →