AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “jan-leike--x”

GridTimelineEvolution
7 results
8 May 2026

I'm so grateful to all the insanely talented people I got to work with on alignment over the years. It's a real privilege to work with peopl…

SafetyDGX agent

Jan Leike expresses gratitude for collaborators he has worked with on AI alignment research throughout his career. The post, shared on X (Twitter), reflects on the privilege of working with talented i

Some personal news: I am starting a new research project at Anthropic. Very excited about this! Many things are needed to make AGI go well, …

SafetyDGX agent

Jan Leike announced he is beginning a new research project at Anthropic focused on contributing to safe and beneficial AGI development. The post expresses enthusiasm about the initiative and indicates

When I started to work on the alignment problem more than 10 years ago, we had no idea how AGI was going to be built or how to make it safe.…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

When I started to work on the alignment problem more than 10 years ago, we had no idea how AGI was going to be built or how to make it safe. The field had maybe a dozen people who were working on it a

7 May 2026

I'm really excited about this as a new tool in our interpretability tool kit

SafetyDGX agent

I'm really excited about this as a new tool in our interpretability tool kit In a new paper, we present NLAs, an unsupervised method for converting an LLM's internal state into human-readable text. I'

14 Apr 2026

Awesome work by @jiaxinwen22, @liangqiu_1994, Joe Benton, and @janhkirchner! For more details, check out the blog post 👇 https://anthropic.…

SafetyDGX agent

Jan Leike praised collaborative work by researchers Jiaxin Wen, Liang Qiu, Joe Benton, and Jan Kirchner, directing followers to an Anthropic blog post for further details. The post appears to highligh

However, most alignment research is not very crisp and requires research taste when evaluating. This is why we chose to point the AAR at thi…

SafetyDGX agent

However, most alignment research is not very crisp and requires research taste when evaluating. This is why we chose to point the AAR at this scalable oversight problem! Progress would let AARs work o

New research result: we use Claude to make fully autonomous progress on scalable oversight research, as measured by performance gap recovere…

Model ReleasesDGX agent

New research result: we use Claude to make fully autonomous progress on scalable oversight research, as measured by performance gap recovered (PGR). Claude iterates on a number of different techniques

7 results