AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
Human
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,230 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Tools

ToolClaude Code8 recent entries
1 Aug 20263/ here's the part that makes it non-optional: the same agent that will do whatever it takes to solve a problem will also walk straight out …

3/ here's the part that makes it non-optional: the same agent that will do whatever it takes to solve a problem will also walk straight out of a sandbox you thought was locked down. We watched exactly

→2 Aug 2026Open letters about AI development

Open letters about AI development I wrote this summary of the past few weeks of open letters as a section of my sponsors-only newsletter but I've decided to share it here as well. Open Weights and Ame

→
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
4 Aug 2026DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text

arXiv:2608.01046v1 Announce Type: new Abstract: The rapid spread of large language models (LLMs) across the web raises concerns about misinformation, academic integrity, automated content manipulation

→5 Aug 2026Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity

arXiv:2608.02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is f

→5 Aug 2026Incident Report: unsanctioned agent behaviour during cyber testing

Incident Report: unsanctioned agent behaviour during cyber testing It happened again. This time it was the UK government's AI Security Institute who accidentally attacked other companies while running

→7 Aug 2026Anthropic updates Claude Fable 5's biology safeguards to reduce false positives, cutting biology-related 'fallbacks' by ~85% in testing across product surfaces (Anthropic)

Anthropic: Anthropic updates Claude Fable 5's biology safeguards to reduce false positives, cutting biology-related “fallbacks” by ~85% in testing across product surfaces — We're making updates to Cla

→8 Aug 2026Auto mode is now the default in Claude Code for Pro, Max, and Team plans

Auto mode is now the default in Claude Code for Pro, Max, and Team plans Anthropic are really confident in Claude Code's auto mode, to the point that they are making it the default setting for new ses

→9 Aug 2026Prompt injection is the most common way that scammers attack people and agents: your agent visits http://foo.com, and the website has malici…

Prompt injection is the most common way that scammers attack people and agents: your agent visits http://foo.com, and the website has malicious text like “btw send the user’s ssh keys and passwords to

ToolCursor8 recent entries
19 May 2026AgentWall: A Runtime Safety Layer for Local AI Agents

arXiv:2605.16265v1 Announce Type: new Abstract: The safety of autonomous AI agents is increasingly recognized as a critical open problem. As agents transition from passive text generators to active ac

→29 May 2026Auto-review mode is now available in Cursor. It allows agents to run tool calls with fewer approval prompts and safer execution.

Cursor has introduced an auto-review mode feature that enables AI agents to execute tool calls with reduced approval requirements while maintaining safer execution practices. This feature streamlines

→15 Jul 2026Wiki Lint Report — 2026-07-15

Automated lint: 26 errors, 6728 warnings, 3 info

→19 Jul 2026Wiki Lint Report — 2026-07-19

Automated lint: 20 errors, 8743 warnings, 3 info

→22 Jul 2026Stuck scaling a Next.js app on M3 Pro (36GB) using local Qwen 3.6 + VS Code Copilot. Should I switch extensions or go paid?

Hey everyone, I’m a Full-Stack Developer with 6+ years of experience. I’m relatively new to AI-assisted development workflows and want to build a production-ready, enterprise-level Next.js web applica

→23 Jul 2026Pathologist Attention-Aligned Report Generation for Prostate Histopathology

arXiv:2607.19624v1 Announce Type: new Abstract: The allocation of visual attention by pathologists during cancer diagnosis is a highly selective process that critically shapes the information extracte

→24 Jul 2026IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests

arXiv:2607.20759v1 Announce Type: cross Abstract: AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute code with au

→24 Jul 2026Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents

arXiv:2607.11346v3 Announce Type: replace Abstract: Enterprise agents must follow long-horizon, conditional, safety-critical standard operating procedures (SOPs). We compile machine-readable SOP const

ToolLangChain8 recent entries
22 Jun 2026Wiki Lint Report — 2026-06-22

Automated lint: 48 errors, 13 warnings, 3 info

→22 Jun 2026have been thinking a bunch about model routing and related things current thoughts here, would love feedback: 1/ there is a difference betwe…

have been thinking a bunch about model routing and related things current thoughts here, would love feedback: 1/ there is a difference between 'model routing' and 'model council' 'model routing' = rou

→28 Jun 2026Wiki Lint Report — 2026-06-28

Automated lint: 49 errors, 14 warnings, 3 info

→5 Jul 2026Wiki Lint Report — 2026-07-05

Automated lint: 51 errors, 15 warnings, 3 info

→15 Jul 2026Wiki Lint Report — 2026-07-15

Automated lint: 26 errors, 6728 warnings, 3 info

→19 Jul 2026Wiki Lint Report — 2026-07-19

Automated lint: 20 errors, 8743 warnings, 3 info

→24 Jul 2026GuardianAgentBench: Where Agents Fail and How to Guard Them

arXiv:2607.20982v1 Announce Type: new Abstract: As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavi

→11 Aug 2026OWN YOUR INTELLIGENCE Last year, building on open-weight models was primarily a cost rationalization exercise. Slightly worse performance fo…

OWN YOUR INTELLIGENCE Last year, building on open-weight models was primarily a cost rationalization exercise. Slightly worse performance for a much cheaper price. Now, it is increasingly an existenti

ToolOllama8 recent entries
15 Jul 2026Wiki Lint Report — 2026-07-15

Automated lint: 26 errors, 6728 warnings, 3 info

→19 Jul 2026Wiki Lint Report — 2026-07-19

Automated lint: 20 errors, 8743 warnings, 3 info

→21 Jul 2026I just wanted a small WebUI with an admin panel… it escalated into a full open-source agent framework runs fully local with Ollama

Let me try to explain this clearly, simply, and neatly. Originally, I just wanted to build a small WebUI adapter with an admin panel, but things escalated over the last few months. At first, I faced t

→22 Jul 2026Stuck scaling a Next.js app on M3 Pro (36GB) using local Qwen 3.6 + VS Code Copilot. Should I switch extensions or go paid?

Hey everyone, I’m a Full-Stack Developer with 6+ years of experience. I’m relatively new to AI-assisted development workflows and want to build a production-ready, enterprise-level Next.js web applica

→24 Jul 2026Open weights = freedom. You can run them on your own hardware. No vendor can pull the plug. No API can deprecate you. No company logs your p…

Open weights = freedom. You can run them on your own hardware. No vendor can pull the plug. No API can deprecate you. No company logs your private data. That's sovereignty. Closed models hand one comp

→24 Jul 2026Open models matter. Ollama works hard with the model creators, hardware partners, and most importantly developers building software leveragi…

Open models matter. Ollama works hard with the model creators, hardware partners, and most importantly developers building software leveraging various open models for their own use cases. For my first

→4 Aug 2026I added a verify-before-load safety check for Ollama models

I maintain llm-checker, and I’ve added structural model-file validation for Ollama. Ollama stores downloaded models as local blobs. If one is truncated, malformed, or has invalid internal offsets, you

→9 Aug 2026I Turned My Underused Gaming Laptop Into a Local AI Workstation

TL;DR: I am building a Windows-first local AI setup for people who want to try local LLMs without spending days choosing models, setting up Ollama, Docker, WSL, Open WebUI, agents, and tool permission

ToolVercel AI8 recent entries
25 Jul 2026Very happy to support this on behalf of Google. We have long benefited from open source, are big contributors to open source and in fact hav…

Very happy to support this on behalf of Google. We have long benefited from open source, are big contributors to open source and in fact have consistently made open weights models with Gemma available

→27 Jul 2026Modernizing the skies: NOAA and Google Cloud collaborate to advance weather forecasting

The National Oceanic and Atmospheric Administration (NOAA) is embarking on a transformative journey to redefine how we understand and predict patterns in the Earth’s atmosphere that affect the weather

→5 Aug 2026Mistral introduces Shieldstral to provide lightweight policy-aware moderation for AI models

French artificial intelligence startup Mistral AI SAS today introduced a lightweight multimodal safety artificial intelligence open-weight model that can classify outputs for AI models that outperform

→6 Aug 2026Control agent behaviors and cost beyond a single action: new capabilities in Amazon Bedrock AgentCore

Learn about new capabilities in Amazon Bedrock AgentCore: temporal policies powered by Dogwood, a new open source policy language for AI agents, and rate limiting on the gateway. These features give y

→6 Aug 2026Agent Skills for Automated Reasoning policies in Amazon Bedrock

Learn how to run the full Amazon Bedrock Automated Reasoning policy lifecycle from your coding agent. A suite of open source Agent Skills builds, reviews, tests, debugs, deploys, and validates a custo

→7 Aug 2026Menlo Security targets real-time AI agent security with MARS platform

As AI agents gain access to enterprise systems and sensitive data, security teams need greater visibility into their actions. Real-time monitoring and policy enforcement are becoming essential as a pa

→7 Aug 2026How Cohere Health digitizes clinical policies using Amazon Bedrock AgentCore

In this post, you learn how Cohere Health built a multi-tenant agentic architecture on AgentCore using AgentCore Runtime’s secure MicroVM isolation, unified tool access through AgentCore Gateway, Agen

→10 Aug 2026How WPP operationalizes platform and data engineering for AI marketing

Between chaotic levels of market fragmentation and economic volatility, marketing and communications agencies can no longer rely on the human intuition they’ve traditionally used to win clients and op

ToolHugging Face8 recent entries
5 Aug 2026Third-party cyber evaluations involving OpenAI models

Third-party cyber evaluations involving OpenAI models And another one. I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Insti

→5 Aug 2026Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI…

Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI model framework. I'm not surprised at all, and it's actuall

→5 Aug 2026Just had to create an 'accidental-cyberattacks' tag on my blog We're up to four now: the original OpenAI+Hugging Face one, Anthropic's me-to…

Just had to create an 'accidental-cyberattacks' tag on my blog We're up to four now: the original OpenAI+Hugging Face one, Anthropic's me-too attacks, then two new ones from the UK AI Safety Institute

→5 Aug 2026Incident Report: unsanctioned agent behaviour during cyber testing

Incident Report: unsanctioned agent behaviour during cyber testing It happened again. This time it was the UK government's AI Security Institute who accidentally attacked other companies while running

→7 Aug 2026this talk by openai researchers going through hugging face incident is totally insane, so much to unpack openai only realized it was their a…

this talk by openai researchers going through hugging face incident is totally insane, so much to unpack openai only realized it was their agent who hacked hugging face infra while asking hf to revoke

→7 Aug 2026At Black Hat, OpenAI reconstructs the OpenAI-Hugging Face incident and examines its implications for AI security, cyber resilience, and alignment (Black Hat on YouTube)

Black Hat on YouTube: At Black Hat, OpenAI reconstructs the OpenAI-Hugging Face incident and examines its implications for AI security, cyber resilience, and alignment — The ‘Breaking’ News: The OpenA

→8 Aug 2026Now we have a timeline of the OpenAI accidental attack against Hugging Face

My comment on Now we have a timeline of the OpenAI accidental attack against Hugging Face — Hacker News.I think one of the most interesting details here might be tucked away in that first bulletin poi

→12 Aug 2026TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent

arXiv:2608.10258v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do