AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “research”

GridTimelineEvolution
25,646 results
5 Aug 2026

MDArena: Evaluating Coding Agents on Realistic Molecular Dynamics Workflows

Model ReleasesDGX agent

arXiv:2608.02642v1 Announce Type: cross Abstract: Accelerating scientific discovery is among the most consequential applications of AI, and computational biomolecular simulation stands out as a partic

4 Aug 2026

ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification

Model ReleasesDGX agent

arXiv:2608.01291v1 Announce Type: new Abstract: We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syrian, E

CRISP: Critical Step Perception for Training Efficient Deep Search Agents

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety
DGX agent

arXiv:2608.01867v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly extended into deep search agents that solve complex questions through multi-step interaction with external

Future Mode Part 2: The foundation for securing agentic browsing

Model ReleasesDGX agent

Editor's Note: Our Future Mode series will give businesses insight into how Chrome Enterprise is approaching AI in the browser. Stay tuned for more blogs in this series.Future Mode Part 2: The foundat

NVIDIA Alpamayo 2 Super, the Frontier Open Model for Robotaxis and Autonomous Vehicles, Now Available for Commercial Use

HardwareDGX agent

For robotaxis and other autonomous vehicles (AVs), the hardest problems aren’t the everyday scenarios. They’re the rare, complex situations that are difficult to anticipate and train for. Handling the

Running fast is not enough, you need fast AND correct An excellent addition from @ArtificialAnlys to make sure that the flashy speed numbers…

Model ReleasesDGX agent

Running fast is not enough, you need fast AND correct An excellent addition from @ArtificialAnlys to make sure that the flashy speed numbers are backed by 100% matching accuracy Announcing the Artific

Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

Model ReleasesDGX agent

arXiv:2608.01666v1 Announce Type: new Abstract: However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open q

3 Aug 2026

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

Model ReleasesDGX agent

arXiv:2607.29626v1 Announce Type: new Abstract: As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly i

Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications

Model ReleasesDGX agent

arXiv:2607.28840v1 Announce Type: new Abstract: Large language models are increasingly deployed in financial applications that combine retrieval, proprietary data, tool use, orchestration logic, monit

CBCT-IQ: A Publicly Available Annotated Cone-Beam CT Dataset for Image Quality Assessment and Benchmarking

Model ReleasesDGX agent

arXiv:2607.29253v1 Announce Type: cross Abstract: Medical image quality plays a critical role in diagnostic accuracy, especially in X-ray-based imaging modalities such as cone-beam computed tomography

M3MAD-Bench: Multi-Dimensional Evaluation of Multi-Agent Debate Across Domains and Modalities

Model ReleasesDGX agent

arXiv:2601.02854v2 Announce Type: replace Abstract: As an agent-level reasoning and coordination paradigm, Multi-Agent Debate (MAD) orchestrates multiple agents through structured debate to improve an

2 Aug 2026

A profile of Jacob Tsimerman, who won the Fields Medal last week and is taking a leave from the University of Toronto to join OpenAI and work on AI safety (Ben Cohen/Wall Street Journal)

SafetyDGX agent

Ben Cohen / Wall Street Journal: A profile of Jacob Tsimerman, who won the Fields Medal last week and is taking a leave from the University of Toronto to join OpenAI and work on AI safety — Jacob Tsim

Open letters about AI development

Model ReleasesDGX agent

Open letters about AI development I wrote this summary of the past few weeks of open letters as a section of my sponsors-only newsletter but I've decided to share it here as well. Open Weights and Ame

Watching @ClementDelangue on @FaceTheNation discussing agentic hacking. “Preventing releases does not work; concentrating behind closed door…

HardwareDGX agent

Watching @ClementDelangue on @FaceTheNation discussing agentic hacking. “Preventing releases does not work; concentrating behind closed doors in just a few organizations doesnt work. What worked in th

1 Aug 2026

If you maintain an AGENTS.md or a CLAUDE.md, this is worth a read. (bookmark it) 288 gold-test evaluated runs across Claude Code and Codex, …

Model ReleasesDGX agent

If you maintain an AGENTS.md or a CLAUDE.md, this is worth a read. (bookmark it) 288 gold-test evaluated runs across Claude Code and Codex, 17 real tasks from 3 repositories, with context-injection st

Ten advances in mathematics and theoretical computer science

Model ReleasesDGX agent

Ten advances in mathematics and theoretical computer science A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending 100,000 on tokens and with

31 Jul 2026

A Distributed Acoustic Sensing Dataset for Vessel Detection and Localization in Submarine Cable Protection

Model ReleasesDGX agent

arXiv:2607.28306v1 Announce Type: cross Abstract: Recent incidents of accidental damage and suspected sabotage to submarine telecommunication and power cables, particularly in the Baltic Sea, have und

AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

Model ReleasesDGX agent

arXiv:2607.28618v1 Announce Type: new Abstract: Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet existing literature-search systems pr

Cloud CISO Perspectives: Why AI Threat Defense is the new boardroom baseline

Local AiDGX agent

Welcome to the second Cloud CISO Perspectives for July 2026. Today, Chris Betz, CISO, Google Cloud, and Alicja Cade, Senior Director, Office of the CISO, Google Cloud, explain what boards of directors

HealthCAT: An Interpretable Encoder-only Transformer Framework for Health Indicator Prediction and Temporal Interpretation of Wearable Sensor Data

ApplicationsDGX agent

arXiv:2607.27635v1 Announce Type: cross Abstract: Wearable sensors continuously capture fine-grained multivariate time-series data, providing opportunities to model behavioural patterns associated wit

IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations

Model ReleasesDGX agent

arXiv:2607.26075v1 Announce Type: cross Abstract: We present IDP AutoOpt, an autonomous LLM agent that discovers high-performing configurations for intelligent document processing (IDP) pipelines. Tun

Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges

AgentsDGX agent

arXiv:2607.26212v1 Announce Type: cross Abstract: Multi-Agent Debate (MAD) is a promising paradigm for improving the accuracy and robustness of Large Language Model (LLM)-based agentic systems. It ena

Oxide and Friends: The Open Weight Revolution with Simon Willison

Model ReleasesDGX agent

Oxide and Friends: The Open Weight Revolution with Simon Willison On Monday Bryan Cantrill and Adam Leventhal invited me to join their podcast to talk about the wild week we've had - with Kimi K3 show

Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork

AgentsDGX agent

arXiv:2607.27177v1 Announce Type: new Abstract: Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume t

The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making

TutorialsDGX agent

arXiv:2607.27179v1 Announce Type: cross Abstract: Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communication among

When Should AI Follow? Task Structure and Joint Adaptation by Human and AI Agents

Local AiDGX agent

arXiv:2504.20903v4 Announce Type: replace-cross Abstract: How should organizations divide and sequence decision tasks between human and artificial agents? We develop a computational model of joint seq

30 Jul 2026

Four Ways to Deploy More Secure AI Agents

HardwareDGX agent

NVIDIA’s AI Red Team found common failure modes in enterprise AI agents: weak access controls, unrestricted code execution via tools, unprotected network egress, and exposure of plaintext secrets. The

Neural Radiance Fields for the Real World: A Survey

ApplicationsDGX agent

arXiv:2501.13104v3 Announce Type: replace Abstract: Neural Radiance Fields (NeRFs) have remodeled 3D scene representation since release. NeRFs can effectively reconstruct complex 3D scenes from 2D ima

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.26873v1 Announce Type: new Abstract: Test-time reinforcement learning (TTRL) enables language models to self-evolve at inference time without labeled feedback. Existing methods rely on answ

29 Jul 2026

DeepSeek V4 Flash isn't just for inference anymore. Fine-tune it on Fireworks with supervised fine-tuning, preference tuning, and combined p…

Model ReleasesDGX agent

DeepSeek V4 Flash isn't just for inference anymore. Fine-tune it on Fireworks with supervised fine-tuning, preference tuning, and combined preference optimization from the managed UI. Reinforcement le

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications

Model ReleasesDGX agent

arXiv:2607.25642v1 Announce Type: cross Abstract: Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models

Thrilled to have @gabepereyra speak at our @sequoia event tmrw on OWN YOUR AI: how to build your own Lab as an application company. Also fea…

Model ReleasesDGX agent

Thrilled to have @gabepereyra speak at our @sequoia event tmrw on OWN YOUR AI: how to build your own Lab as an application company. Also featuring @FireworksAI_HQ @mercor_ai @LangChain @trajectorylabs

28 Jul 2026

Agentic Autoresearch for CT Reconstruction

AgentsDGX agent

arXiv:2607.22824v1 Announce Type: cross Abstract: Comparing CT reconstruction methods fairly is labor-intensive and largely manual, and many benchmarks use idealized data. We ask whether a large langu

Algorithmic Blindness in Large Language Models: A Calibration Study of Performance Prediction

Model ReleasesDGX agent

arXiv:2602.21947v5 Announce Type: replace Abstract: Large language models (LLMs) demonstrate remarkable breadth of knowledge, yet their ability to reason about computational processes remains poorly u

AutoMat: Enabling Automated Crystal Structure Reconstruction from Microscopy via Agentic Tool Use

Model ReleasesDGX agent

arXiv:2505.12650v2 Announce Type: replace-cross Abstract: Reconstructing atomistic crystal structures from a single noisy STEM projection is an ill-posed inverse problem: multiple lattices can explain

Energy Constrained Hierarchical Underwater Monitoring via Local Multi-Agent RAG

HardwareDGX agent

arXiv:2607.24313v1 Announce Type: cross Abstract: Marine life monitoring is limited by strict energy constraints, poor underwater connectivity, and the high cost of transmitting raw multimodal data fr

Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers

AgentsDGX agent

arXiv:2607.24419v1 Announce Type: new Abstract: Deep models have substantially advanced 12-lead ECG classification, yet their refinement still relies heavily on human experts to inspect failures and i

GaitFace: A Multimodal Dataset for Long-Range Person Identification

Model ReleasesDGX agent

arXiv:2607.23542v1 Announce Type: new Abstract: Efficient border control is becoming a significant global challenge, mainly due to severe congestion and extended passenger waiting times. To mitigate t

Neuromorphic Object Detection: An In-Depth Study and Future Directions

Model ReleasesDGX agent

arXiv:2607.23576v1 Announce Type: new Abstract: Conventional frame-based cameras face significant challenges in detecting objects under high-speed motion blur or in low-light environments. Neuromorphi

Novel Claim or Deja Vu? Rethinking 'Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking

Model ReleasesDGX agent

arXiv:2607.23514v1 Announce Type: cross Abstract: Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks

Now, this: 1,100 current/former frontier-AI employees sign a petition calling for US gov't to step in for 'pacing' frontier development

Local AiDGX agent

So, it appears that this is the week of open letters in AI🥲... an open letter signed by current and former employees of OpenAI, Anthropic and Google primarily - calling for a slow-down in frontier AI

Physical AI Governance: From Theory to Practice Across Life Cycle

SafetyDGX agent

arXiv:2607.22877v1 Announce Type: new Abstract: With the emergence of Physical AI, artificial intelligence is extending beyond screen-based applications to embodied systems that perceive, interact wit

Robustness and Cybersecurity in the EU Artificial Intelligence Act

SafetyDGX agent

arXiv:2502.16184v3 Announce Type: replace Abstract: The EU Artificial Intelligence Act (AIA) establishes different legal principles for different types of AI systems. While prior work has sought to cl

Sign-Symmetry Learning Rules are Robust Fine-Tuners

SafetyDGX agent

arXiv:2502.05925v2 Announce Type: replace-cross Abstract: Backpropagation (BP) has long been the predominant method for training neural networks due to its effectiveness. However, numerous alternative

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.24054v1 Announce Type: new Abstract: A correct answer can conceal why an agent succeeded. Once agents change their information state during evaluation, correctness no longer distinguishes i

UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents

Model ReleasesDGX agent

arXiv:2508.00288v5 Announce Type: replace-cross Abstract: Aerial navigation is a fundamental yet underexplored capability in embodied intelligence, enabling agents to operate in large-scale, unstructu

27 Jul 2026

Enigma raises $71M to develop foundation models for robots

IndustryDGX agent

Engima Ltd., a provider of artificial intelligence software for robots, launched today with 71 million in funding. Index Ventures and Ribbit Capital jointly led the seed round with participation from

Last week, we made ChatGPT Work available to enterprises. It's an awesome product. So for ChatGPT Enterprise customers, enroll by August 21,…

ApplicationsDGX agent

Last week, we made ChatGPT Work available to enterprises. It's an awesome product. So for ChatGPT Enterprise customers, enroll by August 21, and each teammate who tries Work for the first time within

Modernizing the skies: NOAA and Google Cloud collaborate to advance weather forecasting

Model ReleasesDGX agent

The National Oceanic and Atmospheric Administration (NOAA) is embarking on a transformative journey to redefine how we understand and predict patterns in the Earth’s atmosphere that affect the weather

Nvidia, Microsoft launch open AI security alliance — without OpenAI, Google, or Anthropic

SafetyDGX agent

Nvidia on Monday said it is joining forces with Microsoft, SpaceX, IBM, and other tech companies to build and share open-source AI security tools. The new Open Secure AI Alliance said open tools are r

ReCowGnition: A Realistic Biometric Benchmark for Cow Face Recognition

Model ReleasesDGX agent

arXiv:2607.22071v1 Announce Type: new Abstract: With the development of precision livestock farming and the advances in computer vision, visual animal biometrics has gained attention. Using biometric

Very cool paper from Microsoft. The idea is to train agents on replayed teacher trajectories instead of live environment rollouts. On-policy…

SafetyDGX agent

Very cool paper from Microsoft. The idea is to train agents on replayed teacher trajectories instead of live environment rollouts. On-policy distillation for agentic tasks is expensive because every u

26 Jul 2026

Has anyone compared pre-training, SFT/LoRA and reinforcement post-training on Qwen3.6-27B?

Model ReleasesDGX agent

Qwen3.6-27B: SFT vs continued pre-training vs RL? I’m interested in adapting Qwen3.6-27B, but I’m increasingly unsure whether conventional SFT/LoRA is the best route if the goal is to add a capability

How US companies flipped from 'tokenmaxxing' to 'thrift-maxxing', mixing cheaper Chinese models with OpenAI and Anthropic, threatening the labs' IPO valuations (Wall Street Journal)

IndustryDGX agent

Wall Street Journal: How US companies flipped from “tokenmaxxing” to “thrift-maxxing”, mixing cheaper Chinese models with OpenAI and Anthropic, threatening the labs' IPO valuations — Companies big and

25 Jul 2026

A profile of Yang Zhilin, founder of Moonshot AI, which faced early doubts over revenue and model capabilities before its Kimi K3 delivered a 'DeepSeek moment' (Financial Times)

Model ReleasesDGX agent

Financial Times: A profile of Yang Zhilin, founder of Moonshot AI, which faced early doubts over revenue and model capabilities before its Kimi K3 delivered a “DeepSeek moment” — Known as ‘Yang the ge

Diffusion LLMs can now handle real agentic work. LLaDA 2.2 is the first large-scale diffusion LLM built to operate as a real agent, planning…

AgentsDGX agent

Diffusion LLMs can now handle real agentic work. LLaDA 2.2 is the first large-scale diffusion LLM built to operate as a real agent, planning, calling tools, and self-correcting across long multi-turn

Sources: DeepSeek told investors it is suspending its second funding round after remarks attributed to Liang Wenfeng on US-China AI competition went viral (Pei Li/Bloomberg)

Model ReleasesDGX agent

Pei Li / Bloomberg: Sources: DeepSeek told investors it is suspending its second funding round after remarks attributed to Liang Wenfeng on US-China AI competition went viral — DeepSeek has told prosp

24 Jul 2026

A joint preliminary evaluation by the UK's AISI and the US' CAISI finds Kimi K3 trails leading US frontier closed weight models on cyber capability (AI Security Institute)

IndustryDGX agent

AI Security Institute: A joint preliminary evaluation by the UK's AISI and the US' CAISI finds Kimi K3 trails leading US frontier closed weight models on cyber capability — The UK Artificial Intellige

Announcing Fugu-Ultra v1.1 🐡 We’ve been thrilled by the reception to the Fugu model family. Thanks to everyone who tried it, shared feedbac…

Model ReleasesDGX agent

Announcing Fugu-Ultra v1.1 🐡 We’ve been thrilled by the reception to the Fugu model family. Thanks to everyone who tried it, shared feedback, and trusted Fugu with real work. Today, we’re releasing Fu

Benchmarking Unlearning for Vision Transformers

Model ReleasesDGX agent

arXiv:2602.20114v2 Announce Type: replace-cross Abstract: Machine unlearning (MU) refers to the post-training capability to remove (the influence of) training examples that are incorrect, biased, or l

← Previous
1…361362363364365…428
Next →