AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “openai”

GridTimelineEvolution
2,581 results
Local Ai

A Mechanistic Explanatory Strategy for XAI

DGX agent

arXiv:2411.01332v5 Announce Type: replace Abstract: Despite significant advancements in XAI, scholars note a persistent lack of solid conceptual foundations and integration with broader scientific dis

local-aiarxiv-cs-lg
23 May 2026
Safety

I once asked a bunch of e/accs how much damage was an acceptable risk relative to take (any) safety precautions. none answered. we may be ab…

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

I once asked a bunch of e/accs how much damage was an acceptable risk relative to take (any) safety precautions. none answered. we may be about to find out. One implication of the below is that we rea

safetygary-marcus--x
23 May 2026
Safety

it ain’t just me who sees the emperor has no clothes

DGX agent

it ain’t just me who sees the emperor has no clothes The AI bubble math doesn't add up. Anthropic spends 3 to make 1 and that’s before you include any and all other costs like staff or electricity. Mi

safetygary-marcus--x
23 May 2026
Safety

Amplifying, Not Learning: Fine-Tuned AI Text Detectors Amplify a Pretrained Direction

DGX agent

arXiv:2605.21653v1 Announce Type: cross Abstract: AI text detectors amplify a pretrained typicality axis; they do not construct an AI-vs-human boundary. On raw encoders before any task supervision, pr

safetyarxiv-cs-cl
22 May 2026
Model Releases

Code Researcher: Deep Research Agent for Large Systems Code and Commit History

DGX agent

arXiv:2506.11060v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based coding agents have shown promising results on coding benchmarks, but their effectiveness on systems code rema

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

I truly miss the age of science and transparency in AI. Especially given how much money and political power and governance and scientific un…

DGX agent

I truly miss the age of science and transparency in AI. Especially given how much money and political power and governance and scientific understanding is at stake. We don’t know for example • How man

model-releasesgary-marcus--x
22 May 2026
Agents

Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling

DGX agent

arXiv:2605.21470v1 Announce Type: new Abstract: Computer-use agents (CUA) automate tasks specified with natural language such as 'order the cheapest item from Taco Bell' by generating sequences of cal

agentsarxiv-cs-lg
21 May 2026
Safety

can’t believe people assume that success on highly verifiable problems in math (where we don’t even know how many tests were performed and h…

DGX agent

can’t believe people assume that success on highly verifiable problems in math (where we don’t even know how many tests were performed and how many might have failed) iautomatically generalize to ever

safetygary-marcus--x
21 May 2026
Model Releases

proof too complicated, Claude help ELI5

DGX agent

proof too complicated, Claude help ELI5 Today, we share a breakthrough on the planar unit distance problem, a famous open question first posed by Paul Erdős in 1946. For nearly 80 years, mathematician

model-releasesjerry-liu--x
21 May 2026
Agents

🚀Qwen3.7-Max just landed at 56.6 on the Artificial Analysis Intelligence Index — a solid 4.8pt jump over Qwen3.6-Max-Preview. @ArtificialAn…

DGX agent

🚀Qwen3.7-Max just landed at 56.6 on the Artificial Analysis Intelligence Index — a solid 4.8pt jump over Qwen3.6-Max-Preview. @ArtificialAnlys ⚡️Sharper sci reasoning, stronger agentic chops, better c

agentsqwen--x
21 May 2026
Industry

It’s a really special time to be alive…some thoughts from training this model 🧵

DGX agent

It’s a really special time to be alive…some thoughts from training this model 🧵 Today, we share a breakthrough on the planar unit distance problem, a famous open question first posed by Paul Erdős in

industrysam-altman--x
20 May 2026
Industry

Once AI starts making solving open problems in novel ways it won’t stop. We are entering the final stage of human solutions to open problems…

DGX agent

Once AI starts making solving open problems in novel ways it won’t stop. We are entering the final stage of human solutions to open problems like this. Feels weird, doesn’t it? Today, we share a break

industryemad-mostaque--x
20 May 2026
Safety

PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents

DGX agent

arXiv:2605.19932v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate over long and recurring external contexts, like document corpora and code repositories. Across in

safetyarxiv-cs-ai
20 May 2026
Safety

the crazy part is that people are “clowning” me without knowing anything about the training or whether anything else other than scaled chang…

DGX agent

the crazy part is that people are “clowning” me without knowing anything about the training or whether anything else other than scaled changed or how the model does on anything else. (or what it costs

safetygary-marcus--x
20 May 2026
Industry

three of the things we are most excited about: 1. AGI accelerating research 2. AGI accelerating companies 3. personal AGI accelerating every…

DGX agent

three of the things we are most excited about: 1. AGI accelerating research 2. AGI accelerating companies 3. personal AGI accelerating everyone in achieving their goals today it was great to announce

industrysam-altman--x
20 May 2026
Model Releases

A Machine With Human-Like Memory Systems

DGX agent

arXiv:2204.01611v3 Announce Type: replace Abstract: Inspired by the cognitive science theory, we explicitly model an agent with both semantic and episodic memory systems, and show that it is better th

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Can LLMs Generate and Solve Linguistic Olympiad Puzzles?

DGX agent

arXiv:2509.21820v2 Announce Type: replace Abstract: In this paper, we introduce a combination of novel and exciting tasks: the solution and generation of linguistic puzzles. We focus on puzzles used i

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Causely: A Causal Intelligence Layer for Enterprise AI A Benchmark Study on SRE and Reliability Workflows

DGX agent

arXiv:2605.18327v1 Announce Type: new Abstract: AI agents deployed into SRE workflows currently derive their understanding of environment state from raw observability telemetry at query time, paying a

model-releasesarxiv-cs-ai
19 May 2026
Hardware

Decart raises $300M for its AI optimization software, world models

DGX agent

Artificial intelligence developer Decart.ai Inc. today announced that it has raised 300 million in funding at a nearly 4 billion valuation. Radical Ventures led the round with participation from Nvidi

hardwaresiliconangle
19 May 2026
Model Releases

Episodic-Semantic Memory Architecture for Long-Horizon Scientific Agents

DGX agent

arXiv:2605.17625v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into persistent scientific collaborators, context window saturation has emerged as a critical bottleneck. Scienti

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps

DGX agent

arXiv:2605.17554v1 Announce Type: new Abstract: Frontier deep research agents (DRAs) plan a research task, synthesize across documents, and return a structured deliverable on demand. They are being de

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

EvilGenie: A Reward Hacking Benchmark

DGX agent

arXiv:2511.21654v2 Announce Type: replace Abstract: We introduce EvilGenie, a benchmark for reward hacking in programming settings. We source problems from LiveCodeBench and create an environment in w

model-releasesarxiv-cs-lg
19 May 2026
Model Releases

Fidelity Probes for Specification--Code Alignment

DGX agent

arXiv:2605.17246v1 Announce Type: cross Abstract: We introduce fidelity probes: natural-language questions generated from a reference artifact with code-derived ground-truth answers, answered from a c

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

The last six months in LLMs in five minutes

DGX agent

I put together these annotated slides from my five minute lightning talk at PyCon US 2026, using the latest iteration of my annotated presentation tool. # I presented this lightning talk at PyCon US 2

model-releasessimon-willison
19 May 2026
Agents

Transfer Learning for Customized Car Racing Environments

DGX agent

arXiv:2605.17928v1 Announce Type: cross Abstract: Transfer Learning, a technique where a model/agent can use the knowledge/expertise that it gained from one task and exploit that to solve another clos

agentsarxiv-cs-lg
19 May 2026
Model Releases

Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classification

DGX agent

arXiv:2605.17691v1 Announce Type: cross Abstract: Automating the classification of negative treatment in legal precedent is a critical yet nuanced NLP task where misclassification carries significant

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale

DGX agent

arXiv:2510.16252v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for web agents demands environments that are both effective for evaluation and efficient enough for large-scale on

model-releasesarxiv-cs-cl
19 May 2026
Safety

When Vision Speaks for Sound

DGX agent

arXiv:2605.16403v1 Announce Type: new Abstract: Despite rapid progress in video-capable MLLMs, we find that their apparent audio understanding in videos is often vision-driven: models rely on visual c

safetyarxiv-cs-cv
19 May 2026
Research

An LLM-RAG Approach for Healthy Eating Index-Informed Personalized Food Recommendations

DGX agent

arXiv:2605.15213v1 Announce Type: cross Abstract: Diet quality is a leading determinant of chronic disease risk. Advances in artificial intelligence (AI) have enabled food recommendation systems to ad

researcharxiv-cs-ai
18 May 2026
Industry

Musk v. Altman: Elon Musk says the judge and jury 'never actually ruled on the merits of the case, just on a calendar technicality' and he will file an appeal (Elon Musk/@elonmusk)

DGX agent

Elon Musk / @elonmusk: Musk v. Altman: Elon Musk says the judge and jury “never actually ruled on the merits of the case, just on a calendar technicality” and he will file an appeal — Regarding the Op

industrytechmeme
18 May 2026
Industry

Musk v. Altman jury unanimously rejects Elon Musk's claims against Sam Altman; the judge dismisses two additional claims citing the statute of limitation (CNBC)

DGX agent

CNBC: Musk v. Altman jury unanimously rejects Elon Musk's claims against Sam Altman; the judge dismisses two additional claims citing the statute of limitation — After less than two hours of deliberat

industrytechmeme
18 May 2026
Tools

My AI Engineer Singapore journey: Most exciting: First time traveling abroad after joining http://Z.ai, and my very first English talk on st…

DGX agent

My AI Engineer Singapore journey: Most exciting: First time traveling abroad after joining http://Z.ai, and my very first English talk on stage. Most challenging: Spoke in a three-speaker session alon

toolsswyx--x
18 May 2026
Applications

So would this be legal under the ruling: Set up a charity, raise $100m etc When setting it up say you may set up a for profit & raise billio…

DGX agent

So would this be legal under the ruling: Set up a charity, raise $100m etc When setting it up say you may set up a for profit & raise billions Three years later set up for profit with IP transferred f

applicationsemad-mostaque--x
18 May 2026
Research

Witchcraft, fast local semantic search on top of SQLite [P]

DGX agent

Witchcraft is a Rust reimplementation of Stanford's XTR-Warp semantic search engine that uses a single-file SQLite database for storage, enabling client-side deployment. The system operates completely

researchr-machinelearning
18 May 2026
Safety

What I am about to describe ain’t AGI; it’s a sign of a trillion dollar trainwreck. If I had told you in 2022 that the 2026 version of GPT (…

DGX agent

What I am about to describe ain’t AGI; it’s a sign of a trillion dollar trainwreck. If I had told you in 2022 that the 2026 version of GPT (which by the way would only be GPT 5.5 and not GPT-6 or 7 li

safetygary-marcus--x
17 May 2026
Tools

getting ready for our first Cabinet Minister ever speaking not as a politician, but as a @NanoClaw_AI user and AI Engineer!

DGX agent

getting ready for our first Cabinet Minister ever speaking not as a politician, but as a @NanoClaw_AI user and AI Engineer! All @aiDotEngineer SG talks kick off in 22 mins! Tune in live: https://www.y

toolsswyx--x
16 May 2026
Safety

holy shit lmao @Gavriel_Cohen he's seriously using this thing for conducting the foreign policy/parliamentary affairs of singapore - and sha…

DGX agent

holy shit lmao @Gavriel_Cohen he's seriously using this thing for conducting the foreign policy/parliamentary affairs of singapore - and sharing his stack on how he is hacking around WhatsApp and doin

safetyswyx--x
16 May 2026
Hardware

SF vibes are frenetic over the huge divide in outcomes and career uncertainty for software engineers; over 5 years ~10K people in AI attained retirement wealth (Deedy/@deedydas)

DGX agent

Deedy / @deedydas: SF vibes are frenetic over the huge divide in outcomes and career uncertainty for software engineers; over 5 years ~10K people in AI attained retirement wealth — The vibes in SF fee

hardwaretechmeme
16 May 2026
Tools

Will be delivering my talk on RL and RLMS after 4pm SGT 1am PT. Tune in to the stream if you're up!

DGX agent

Will be delivering my talk on RL and RLMS after 4pm SGT 1am PT. Tune in to the stream if you're up! All @aiDotEngineer SG talks kick off in 22 mins! Tune in live: https://www.youtube.com/watch?v=_xQnS

toolsswyx--x
16 May 2026
Model Releases

Agentic Design of Compositional Descriptors via Autoresearch for Materials Science Applications

DGX agent

arXiv:2605.14671v1 Announce Type: cross Abstract: Autoresearch offers a flexible paradigm for automating scientific tasks, in which an AI agent proposes, implements, evaluates, and refines candidate s

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

MultiEmo-Bench: Multi-label Visual Emotion Analysis for Multi-modal Large Language Models

DGX agent

arXiv:2605.14635v1 Announce Type: cross Abstract: This paper introduces a multi-label visual emotion analysis benchmark dataset for comprehensively evaluating the ability of multimodal large language

model-releasesarxiv-cs-ai
15 May 2026
Safety

Progent: Securing AI Agents with Privilege Control

DGX agent

arXiv:2504.11703v3 Announce Type: replace-cross Abstract: AI agents interact with external environments through tool calls, exposing them to attacks like indirect prompt injection that can trigger una

safetyarxiv-cs-ai
15 May 2026
Industry

Behold, the Elon Musk jackass trophy

DGX agent

Yesterday, in Musk v. Altman, before the jurors came in, Sam Altman's team passed up what looked - from a distance - like a little league trophy. It was not. Yvonne Gonzalez Rogers had the lawyers rea

industrythe-verge-ai
14 May 2026
Model Releases

From Pixels to BFS: High Maze Accuracy Does Not Imply Visual Planning

DGX agent

arXiv:2603.26839v2 Announce Type: replace-cross Abstract: How do multimodal models solve visual spatial tasks -- through genuine planning, or through brute-force search in token space? We introduce ex

model-releasesarxiv-cs-cv
14 May 2026
Research

Just Ask for a Table: A Thirty-Token User Prompt Defeats Sponsored Recommendations in Twelve LLMs

DGX agent

arXiv:2605.12772v1 Announce Type: new Abstract: Wu et al. (2026) showed that most frontier large language models (LLMs) recommend a sponsored, roughly twice-as-expensive flight when their system promp

researcharxiv-cs-cv
14 May 2026
Industry

Blue Origin may need external funding to hit ambitious launch targets

DGX agent

Blue Origin CEO Dave Limp told employees that the company needs external capital to significantly increase launch frequency, stating that achieving established launch goals requires 'a large amount of

industryars-technica
13 May 2026
Model Releases

ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?

DGX agent

arXiv:2605.11086v1 Announce Type: cross Abstract: AI agents are rapidly gaining capabilities that could significantly reshape cybersecurity, making rigorous evaluation urgent. A critical capability is

model-releasesarxiv-cs-lg
13 May 2026
Agents

Great example of why you should 1. Run your agent on a separate machine from the sandbox it uses (e.g. sandbox as a tool) 2. Never set env v…

DGX agent

Great example of why you should 1. Run your agent on a separate machine from the sandbox it uses (e.g. sandbox as a tool) 2. Never set env vars in your sandbox. Instead, use something like LangSmith’s

agentsharrison-chase--x
13 May 2026
← Previous
1…4647484950…54
Next →