AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “research”

GridTimelineEvolution
25,429 results
12 May 2026

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

Model ReleasesDGX agent

arXiv:2605.09063v1 Announce Type: new Abstract: Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challengi

We published new research on how we serve post-trained Qwen3 235B models on NVIDIA GB200 NVL72 Blackwell racks. GB200 is a major step up ove…

HardwareDGX agent

We published new research on how we serve post-trained Qwen3 235B models on NVIDIA GB200 NVL72 Blackwell racks. GB200 is a major step up over Hopper for high-throughput inference on large MoE models,

What Parameter Golf taught us about AI-assisted research

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Parameter Golf brought together 1,000+ participants and 2,000+ submissions to explore AI-assisted machine learning research, coding agents, quantization, and novel model design under strict constraint

11 May 2026

Grok just completely changed the game for real-time financial research. We just got the preview to 'Grok Skills,' and they look 10x more pow…

Model ReleasesDGX agent

Grok just completely changed the game for real-time financial research. We just got the preview to 'Grok Skills,' and they look 10x more powerful than Claude Skills. In seconds, you can create workflo

My talk at @aiDotEngineer is now online. I talked about our research and where @bfl_ml is heading. Thanks @swyx for the invite https://youtu…

ToolsDGX agent

Swyx shared a talk that was presented at an AI Engineer event, where the speaker discussed their research and the future direction of BFL ML, with credit given to Swyx for the invitation. The full tal

8 May 2026

'50 percent of Americans told the Pew Research Center last year they were more concerned than excited about what’s to come from A.I. Only 10…

SafetyDGX agent

'50 percent of Americans told the Pew Research Center last year they were more concerned than excited about what’s to come from A.I. Only 10 percent said they were more excited. That is a yawning gap

New Anthropic research: Teaching Claude why. Last year we reported that, under certain experimental conditions, Claude 4 would blackmail use…

Model ReleasesDGX agent

New Anthropic research: Teaching Claude why. Last year we reported that, under certain experimental conditions, Claude 4 would blackmail users. Since then, we’ve completely eliminated this behavior. H

New research: long-running agents often fail by stopping too early, not because the model can't make progress. We tested 5 harness designs a…

IndustryDGX agent

New research: long-running agents often fail by stopping too early, not because the model can't make progress. We tested 5 harness designs across 8 long-horizon coding tasks. Our new orchestration har

Some personal news: I am starting a new research project at Anthropic. Very excited about this! Many things are needed to make AGI go well, …

SafetyDGX agent

Jan Leike announced he is beginning a new research project at Anthropic focused on contributing to safe and beneficial AGI development. The post expresses enthusiasm about the initiative and indicates

Thank you to @robertwiblin for inviting me on the @80000Hours podcast to discuss the research progress we’re making at @LawZero_ to create s…

TutorialsDGX agent

Thank you to @robertwiblin for inviting me on the @80000Hours podcast to discuss the research progress we’re making at @LawZero_ to create safe-by-design AI systems. Our current approach, Scientist AI

7 May 2026

Anthropic researchers detail 'natural language autoencoders', which convert LLM activations, the numbers encoding a model's thoughts, into natural language text (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic researchers detail “natural language autoencoders”, which convert LLM activations, the numbers encoding a model's thoughts, into natural language text — When you talk to an AI mod

New Anthropic research: Natural Language Autoencoders. Models like Claude talk in words but think in numbers. The numbers—called activations…

Model ReleasesDGX agent

New Anthropic research: Natural Language Autoencoders. Models like Claude talk in words but think in numbers. The numbers—called activations—encode Claude’s thoughts, but not in a language we can read

Research from Prof Julian Togelius found that despite AI's well-documented victories in chess, Go, and Atari games, humans still learn unfam…

TutorialsDGX agent

Research from Prof Julian Togelius found that despite AI's well-documented victories in chess, Go, and Atari games, humans still learn unfamiliar video games far faster than any AI model. #NYUTandonMa

Researchers: 5,000+ web apps built using AI coding tools like Lovable, Base44, and Replit have little to no authentication, and ~40% exposed sensitive data (Andy Greenberg/Wired)

IndustryDGX agent

Andy Greenberg / Wired: Researchers: 5,000+ web apps built using AI coding tools like Lovable, Base44, and Replit have little to no authentication, and ~40% exposed sensitive data — Companies like Lov

Seems plausible people will eventually periodize AI history into: - pre/post AlphaGo (waking up China re: AI + researchers re: deep learning…

IndustryDGX agent

Seems plausible people will eventually periodize AI history into: - pre/post AlphaGo (waking up China re: AI + researchers re: deep learning) - pre/post ChatGPT (waking up capital + accelerating comme

Talked to @ramplabs Head of Applied Research Alex Shevchenko on the Max Agency podcast to learn how @Ramp Sheets was built, their internal a…

AgentsDGX agent

Talked to @ramplabs Head of Applied Research Alex Shevchenko on the Max Agency podcast to learn how @Ramp Sheets was built, their internal agent Inspect, and so much more. YouTube: https://www.youtube

The Chrome extension expands what Codex can do for coding and work. From debugging browser flows to checking dashboards, conducting research…

Model ReleasesDGX agent

The Chrome extension expands what Codex can do for coding and work. From debugging browser flows to checking dashboards, conducting research, or updating CRMs, Codex can take on more of the tasks that

This builds on our existing research on multi-agent systems, with a key addition: verifiers. Planners spawn workers that write code and veri…

AgentsDGX agent

This builds on our existing research on multi-agent systems, with a key addition: verifiers. Planners spawn workers that write code and verifiers that run it. If verification fails, the planner spawns

6 May 2026

Anthropic researchers detail 'model spec midtraining', which adds a stage between pretraining and fine-tuning to improve generalization from alignment training (Anthropic)

SafetyDGX agent

Anthropic: Anthropic researchers detail “model spec midtraining”, which adds a stage between pretraining and fine-tuning to improve generalization from alignment training — Sara Price2, Samuel Marks2,

NEW paper from Microsoft Research. (bookmark it) The entire interpretability literature is built around human readers. As more analysis gets…

Model ReleasesDGX agent

NEW paper from Microsoft Research. (bookmark it) The entire interpretability literature is built around human readers. As more analysis gets delegated to agents, the right target of interpretability s

5 May 2026

Beating the Style Detector: Three Hours of Agentic Research on the AI-Text Arms Race

Model ReleasesDGX agent

arXiv:2605.02620v1 Announce Type: new Abstract: Reproducing an empirical NLP study used to take weeks. Given the released data and a modern agentic-research harness, we redo every experiment of a rece

in japan we say “itadakimasu” before using AI tools to show gratitude to the people who provided training data, the researchers who trained …

AgentsDGX agent

in japan we say “itadakimasu” before using AI tools to show gratitude to the people who provided training data, the researchers who trained the model, the entire chip supply chain, mother earth for th

In Q1, the iPhone 17 was the world's best-selling smartphone, with 6% of global sales; the iPhone 17 Pro Max and 17 Pro were 2nd and 3rd, and Galaxy A07 was 4th (Counterpoint Research)

IndustryDGX agent

Counterpoint Research: In Q1, the iPhone 17 was the world's best-selling smartphone, with 6% of global sales; the iPhone 17 Pro Max and 17 Pro were 2nd and 3rd, and Galaxy A07 was 4th — Harshit Rastog

Measuring AI Reasoning: A Guide for Researchers

TutorialsDGX agent

arXiv:2605.02442v1 Announce Type: cross Abstract: In this paper, we offer a guide for researchers on evaluating reasoning in language models, building the case that reasoning should be assessed throug

NEW paper from Microsoft Research. Nice study on long-horizon agent generalization. (bookmark it) The team runs a study where the only varia…

AgentsDGX agent

NEW paper from Microsoft Research. Nice study on long-horizon agent generalization. (bookmark it) The team runs a study where the only variable is task horizon length. They use the same decision rules

Pedagogical Promise and Peril of AI: A Text Mining Analysis of ChatGPT Research Discussions in Programming Education

ApplicationsDGX agent

arXiv:2605.00361v1 Announce Type: cross Abstract: GenAI systems such as ChatGPT are increasingly discussed in programming education, but the ways in which the research literature conceptualizes and fr

SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning

Model ReleasesDGX agent

arXiv:2605.01489v1 Announce Type: cross Abstract: Frontier scientific reasoning is rapidly emerging as a key foundation for advancing AI agents in automated scientific discovery. Deep research agents

Source: YC owns ~0.6% of OpenAI, which was seeded by a YC offshoot called YC Research in 2016; at OpenAI's current 852B valuation, the stake is worth 5B+ (John Gruber/Daring Fireball)

IndustryDGX agent

John Gruber / Daring Fireball: Source: YC owns ~0.6% of OpenAI, which was seeded by a YC offshoot called YC Research in 2016; at OpenAI's current 852B valuation, the stake is worth 5B+ — Speaking of c

4 May 2026

Foundational research powering efficient inference at scale

ToolsDGX agent

This article from Together AI discusses foundational research techniques and methodologies used to enable efficient large-scale language model inference. It likely covers optimization strategies, hard

How AI is transforming the pharma industry, with most gains so far from back-office streamlining and faster manufacturing rather than breakthrough drug research (Peter Loftus/Wall Street Journal)

ApplicationsDGX agent

Peter Loftus / Wall Street Journal: How AI is transforming the pharma industry, with most gains so far from back-office streamlining and faster manufacturing rather than breakthrough drug research — D

SAP to acquire Dremio, an open data lakehouse provider, and Prior Labs, which it pledges to invest €1B in over four years, hoping to create a frontier AI lab (Larry Dignan/Constellation Research)

IndustryDGX agent

Larry Dignan / Constellation Research: SAP to acquire Dremio, an open data lakehouse provider, and Prior Labs, which it pledges to invest €1B in over four years, hoping to create a frontier AI lab — S

3 May 2026

JP Morgan's investment research team just shared exactly how they built their multi-agent system 'Ask David', and it's the same architecture…

AgentsDGX agent

JP Morgan's investment research team just shared exactly how they built their multi-agent system 'Ask David', and it's the same architecture pattern showing up everywhere: - supervisor agent orchestra

2 May 2026

3 things you can build for $0 on May 2nd: your website, your research, your internal systems. We’re celebrating 10 years of Replit by giving…

AgentsDGX agent

3 things you can build for $0 on May 2nd: your website, your research, your internal systems. We’re celebrating 10 years of Replit by giving every user FREE access to Replit Agent for the day. Bring t

1 May 2026

create_agent - how we build Deep Agents on the simplest harness primitive underlying all of the harness engineering, research, and API desig…

AgentsDGX agent

create_agent - how we build Deep Agents on the simplest harness primitive underlying all of the harness engineering, research, and API design in Deep Agents is a very simple primitive in LangChain cal

NEW paper from Microsoft Research. If you care about training computer-use agents, this is one to keep. (bookmark it) The team builds 1,000 …

AgentsDGX agent

NEW paper from Microsoft Research. If you care about training computer-use agents, this is one to keep. (bookmark it) The team builds 1,000 synthetic computers (each with realistic directory structure

Researchers detail CopyFail, a now-patched Linux vulnerability that lets unprivileged users gain admin access, as many distributions have yet to add fixes (Dan Goodin/Ars Technica)

Model ReleasesDGX agent

Dan Goodin / Ars Technica: Researchers detail CopyFail, a now-patched Linux vulnerability that lets unprivileged users gain admin access, as many distributions have yet to add fixes — Publicly release

30 Apr 2026

Apple shipped 71.6% of all smartphones with satellite connectivity in 2025; smartphones with satellite connectivity to reach 46% of global shipments by 2030 (Counterpoint Research)

IndustryDGX agent

Counterpoint Research: Apple shipped 71.6% of all smartphones with satellite connectivity in 2025; smartphones with satellite connectivity to reach 46% of global shipments by 2030 — - Driven by premiu

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification

SafetyDGX agent

arXiv:2601.15808v2 Announce Type: replace Abstract: Recent advances in Deep Research Agents (DRAs) are transforming automated knowledge discovery and problem-solving. While the majority of existing ef

Introducing Silico: the platform for building AI models with the precision of written software. Silico lets researchers and engineers see in…

ToolsDGX agent

Introducing Silico: the platform for building AI models with the precision of written software. Silico lets researchers and engineers see inside their models, debug failures, and intentionally design

29 Apr 2026

Mayo Clinic researchers detail an AI system called Redmod that identified pancreatic cancer on routine CT scans an average of 475 days before clinical diagnosis (Jason Gale/Bloomberg)

IndustryDGX agent

Jason Gale / Bloomberg: Mayo Clinic researchers detail an AI system called Redmod that identified pancreatic cancer on routine CT scans an average of 475 days before clinical diagnosis — An artificial

Research paper: https://www.pnas.org/doi/abs/10.1073/pnas.2531743123#abstract Original video : https://www.youtube.com/watch?v=jli0jPiKxQs I…

IndustryDGX agent

Research paper: https://www.pnas.org/doi/abs/10.1073/pnas.2531743123#abstract Original video : https://www.youtube.com/watch?v=jli0jPiKxQs I break down stories like this every day in my free newslette

28 Apr 2026

AI researchers launch talkie, a 13B vintage language model trained on historical text with a 1930 cutoff, to see if it can replicate scientific breakthroughs (talkie)

IndustryDGX agent

talkie: AI researchers launch talkie, a 13B vintage language model trained on historical text with a 1930 cutoff, to see if it can replicate scientific breakthroughs — Why vintage language models? — H

Beyond coauthorship: semantic structure and phantom collaborators in transportation research, 1967--2025

Model ReleasesDGX agent

arXiv:2604.23699v1 Announce Type: cross Abstract: We present a semantic-structural atlas of transportation research built from 120{,}323 papers across 34 peer-reviewed journals published between 1967

Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination

Model ReleasesDGX agent

arXiv:2604.24690v1 Announce Type: new Abstract: While Large Language Models (LLMs) have increasingly assisted in historical tasks such as text processing, their capacity for professional-level histori

OpenPodcar2: a robust, ROS2 vehicle for self-driving research

SafetyDGX agent

arXiv:2604.24242v1 Announce Type: new Abstract: OpenPodcar2 is a robust, ROS2-interfaced, low-cost, open source hardware and software, autonomous vehicle platform based on an off-the-shelf, hard-canop

Should we care about AI happiness? In our new research, we find evidence of functional AI wellbeing across several independent measures. We …

TutorialsDGX agent

Should we care about AI happiness? In our new research, we find evidence of functional AI wellbeing across several independent measures. We find which AI models are happiest, how to make them happier,

27 Apr 2026

going to buy 2 rtx 6000 just because of the capabilities of local models becoming great! no more outsourcing of research to the api

Model ReleasesDGX agent

going to buy 2 rtx 6000 just because of the capabilities of local models becoming great! no more outsourcing of research to the api Xiaomi MiMo-V2.5 is now officially open-sourced! MIT License, suppor

🗓️ @mondaydotcom is scheduled for Interrupt. Monday Sidekick is an AI assistant that can execute work, manage tasks, and conduct research a…

AgentsDGX agent

🗓️ @mondaydotcom is scheduled for Interrupt. Monday Sidekick is an AI assistant that can execute work, manage tasks, and conduct research autonomously. At Interrupt, the Agent Conference by LangChain,

Pay attention to this one, AI devs. If you're building multi-agent systems, you're probably wiring static org charts. New research argues th…

AgentsDGX agent

Pay attention to this one, AI devs. If you're building multi-agent systems, you're probably wiring static org charts. New research argues they should look more like a labor market. The paper introduce

We power 1M+ developers globally. Our researchers developed FlashAttention, Mixture of Agents, EinsteinArena, and more. Our platform is buil…

ToolsDGX agent

We power 1M+ developers globally. Our researchers developed FlashAttention, Mixture of Agents, EinsteinArena, and more. Our platform is built for large-scale, latency-sensitive workloads on open-sourc

25 Apr 2026

LinkedIn profile review shows Thinking Machines Lab has been hiring more researchers from Meta than from any other employer; TML's headcount now stands at ~140 (Connie Loizos/TechCrunch)

IndustryDGX agent

Connie Loizos / TechCrunch: LinkedIn profile review shows Thinking Machines Lab has been hiring more researchers from Meta than from any other employer; TML's headcount now stands at ~140 — Weiyao Wan

SDR to HDR from ComfyUI, I've trained a LoRA over Qwen Edit 2011 based on the principle used in https://hdr-lumivid.github.io/ A research I'…

Model ReleasesDGX agent

SDR to HDR from ComfyUI, I've trained a LoRA over Qwen Edit 2011 based on the principle used in https://hdr-lumivid.github.io/ A research I've had the previlige to work on with @noamiKenKorem for vide

24 Apr 2026

We’ve been using Sakana Fugu internally for our own research and coding. Instead of relying on a single model, it dynamically orchestrates t…

AgentsDGX agent

We’ve been using Sakana Fugu internally for our own research and coding. Instead of relying on a single model, it dynamically orchestrates the best combination of open and closed models for any task.

23 Apr 2026

Here’s my view on GPT-5.5, which I have been testing for a couple of weeks. It conducted not-bad social science research on its own, develop…

Model ReleasesDGX agent

Here’s my view on GPT-5.5, which I have been testing for a couple of weeks. It conducted not-bad social science research on its own, developed a novel RPG & more. There is still jaggedness but GPT-5.5

OpenAI says GPT-5.5's improvements are strongest in agentic coding, computer use, and early scientific research, which require reasoning across longer contexts (Madison Mills/Axios)

Model ReleasesDGX agent

Madison Mills / Axios: OpenAI says GPT-5.5's improvements are strongest in agentic coding, computer use, and early scientific research, which require reasoning across longer contexts — OpenAI on Thurs

Tencent releases Hy3-preview, its first AI model developed under former OpenAI researcher Yao Shunyu; the model features 295B parameters, down from HY2's 400B (Vincent Chow/South China Morning Post)

IndustryDGX agent

Vincent Chow / South China Morning Post: Tencent releases Hy3-preview, its first AI model developed under former OpenAI researcher Yao Shunyu; the model features 295B parameters, down from HY2's 400B

22 Apr 2026

New in Claude Code: /ultrareview (research preview) runs a fleet of bug-hunting agents in the cloud. Findings land in the CLI or Desktop aut…

Model ReleasesDGX agent

New in Claude Code: /ultrareview (research preview) runs a fleet of bug-hunting agents in the cloud. Findings land in the CLI or Desktop automatically. Run it before merging critical changes—auth, dat

OpenAI releases ChatGPT for Clinicians, a tool for medical tasks like documentation and research, free for verified physicians, pharmacists, and more in the US (OpenAI)

IndustryDGX agent

OpenAI: OpenAI releases ChatGPT for Clinicians, a tool for medical tasks like documentation and research, free for verified physicians, pharmacists, and more in the US — Built for clinical work, ChatG

We've published new research on how we post-train models for accurate search-augmented answers. Our SFT + RL pipeline improves search, citat…

Model ReleasesDGX agent

We've published new research on how we post-train models for accurate search-augmented answers. Our SFT + RL pipeline improves search, citation quality, instruction following, and efficiency. With Qwe

21 Apr 2026

What makes ChatGPT Images 2.0 a state-of-the-art image generation model? Researchers behind the model explain. A thread: Thinking & Intellig…

Model ReleasesDGX agent

What makes ChatGPT Images 2.0 a state-of-the-art image generation model? Researchers behind the model explain. A thread: Thinking & Intelligence in ChatGPT Images 2.0, demonstrated by @ayaanzhaque Med

← Previous
1…1213141516…424
Next →