AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,555 results
6 May 2026

Hello from Code with Claude!

Model ReleasesDGX agent

This post likely announces or introduces 'Code with Claude,' a tool or service related to Claude AI that enables code generation or programming assistance. The announcement comes from Boris Cherny, a

Hierarchical Memorization in Large Language Models: Evidence from Citation Generation

Model ReleasesDGX agent

arXiv:2511.08877v2 Announce Type: replace Abstract: Large language models (LLMs) generate fluent text across a wide range of tasks, but the fabrication of non-existent academic citations remains a cri

HistCAD: A Constraint-Aware Parametric History-Based CAD Representation, Dataset, and Benchmark with Industrial Complexity

Model ReleasesDGX agent

arXiv:2602.19171v2 Announce Type: replace-cross Abstract: Parametric CAD sequences are reusable because dimensional and geometric constraints govern how parameter changes propagate. Existing CAD gener


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

How Language Models Process Negation

Model ReleasesDGX agent

arXiv:2605.03052v1 Announce Type: new Abstract: We study how Large Language Models (LLMs) process negation mechanistically. First, we establish that even though open-weight models often provide wrong

Hybrid Models for Natural Language Reasoning: The Case of Syllogistic Logic

Model ReleasesDGX agent

arXiv:2510.09472v2 Announce Type: replace Abstract: Despite the remarkable progress in neural models, their ability to generalize, a cornerstone for applications such as logical reasoning, remains a c

I 'also' asked ChatGPT (and Gemini for good measure) how it felt to be an AI.

Model ReleasesDGX agent

This Reddit post documents a user's experiment asking ChatGPT and Gemini about their subjective experience of being an AI, exploring how these language models respond to philosophical questions about

'I Don't Know' -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation

Model ReleasesDGX agent

arXiv:2605.00957v1 Announce Type: cross Abstract: Achieving the right amount of trust in AI systems is important, but challenging. The problem is exacerbated with the rise of Large Language Models (LL

I'm at the Claude w/ Code event in San Francisco, and I'll be live blogging the keynote here https://simonwillison.net/2026/May/6/code-w-cla…

Model ReleasesDGX agent

Simon Willison provided live coverage of the Claude with Code event keynote in San Francisco on May 6, 2026, via his blog. The post documents real-time updates and commentary from the keynote presenta

🧠 Introducing NeuralBench: a unified, open-source framework to benchmark NeuroAI models. v1.0: 36 EEG tasks, 94 datasets, task-specific + f…

Model ReleasesDGX agent

🧠 Introducing NeuralBench: a unified, open-source framework to benchmark NeuroAI models. v1.0: 36 EEG tasks, 94 datasets, task-specific + foundation models. MEG/fMRI ready. MIT-licensed, FAIR's Brain

Introducing the ChatGPT Futures Class of 2026—26 honorees from the first graduating class to have had ChatGPT throughout all four years of u…

Model ReleasesDGX agent

Introducing the ChatGPT Futures Class of 2026—26 honorees from the first graduating class to have had ChatGPT throughout all four years of university, who used AI to: - Map 1.5M previously unknown obj

IRIS: Intent Resolution via Inference-time Saccades for Open-Ended VQA in Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2602.16138v2 Announce Type: replace Abstract: We introduce IRIS (Intent Resolution via Inference-time Saccades), a novel training-free approach that uses eye-tracking data in real-time to resolv

ISAAC: Auditing Causal Reasoning in Deep Models for Drug-Target Interaction

Model ReleasesDGX agent

arXiv:2605.02962v1 Announce Type: new Abstract: Deep learning models for drug--target interaction (DTI) prediction often achieve strong benchmark performance without necessarily relying on mechanistic

It feels like agent harness evolution runs on two axes that usually get conflated. There’s the temporal axis: simplify as models improve, st…

Model ReleasesDGX agent

It feels like agent harness evolution runs on two axes that usually get conflated. There’s the temporal axis: simplify as models improve, stripping components that compensated for limitations the new

It's been an amazing start to Code w/ Claude! Love hearing what people are building with Claude Code and getting feedback on what we can do …

Model ReleasesDGX agent

Boris Cherny shares positive feedback about the early adoption of Claude Code, expressing enthusiasm for the projects users are building with the tool and highlighting the importance of community feed

Jiao: Bridging Isolation and Customization in Mixed Criticality Robotics

Model ReleasesDGX agent

arXiv:2605.03641v1 Announce Type: new Abstract: Consumer robotics demands consolidation of safety-critical control, perception pipelines, and user applications on shared multicore platforms. While sta

Joint Energy Management and Coordinated AIGC Workload Scheduling for Distributed Data Centers: A Diffusion-Aided Reward Shaping Approach

Model ReleasesDGX agent

arXiv:2605.02965v1 Announce Type: new Abstract: Artificial intelligence-generated content (AIGC) has emerged as a transformative paradigm for automating the creation of diverse and customized content,

Kernel Affine Hull Machines for Compute-Efficient Query-Side Semantic Encoding

Model ReleasesDGX agent

arXiv:2605.02950v1 Announce Type: new Abstract: Transformer-based semantic retrieval is highly effective, yet in many deployments the dominant cost lies in online query encoding rather than corpus ind

Label-Efficient School Detection from Aerial Imagery via Weakly Supervised Pretraining and Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.03968v1 Announce Type: new Abstract: Accurate school detection is essential for supporting education initiatives, including infrastructure planning and expanding internet connectivity to un

Learning Dynamics of Zeroth-Order Optimization: A Kernel Perspective

Model ReleasesDGX agent

arXiv:2605.03373v1 Announce Type: new Abstract: Classical optimization theory establishes that zeroth-order (ZO) algorithms suffer from a dimension-dependent slowdown, with convergence rates typically

LightSBB-M: Bridging Schrodinger and Bass for Generative Diffusion Modeling

Model ReleasesDGX agent

arXiv:2601.19312v2 Announce Type: replace Abstract: The Schrodinger Bridge and Bass (SBB) formulation, which jointly controls drift and volatility, is an established extension of the classical Schrodi

LitVISTA: A Benchmark for Narrative Orchestration in Literary Text

Model ReleasesDGX agent

arXiv:2601.06445v2 Announce Type: replace Abstract: Computational narrative analysis aims to capture rhythm, tension, and emotional dynamics in literary texts. Existing large language models can gener

Live blog: Code w/ Claude 2026

Model ReleasesDGX agent

Simon Willison's live blog covers Anthropic's Code w/ Claude 2026 event, documenting the morning keynote sessions with real-time updates. The event featured announcements including updates to Claude m

LiveFMBench: Unveiling the Power and Limits of Agentic Workflows in Specification Generation

Model ReleasesDGX agent

arXiv:2605.01394v1 Announce Type: cross Abstract: Formal specification is essential for rigorous program verification, yet writing correct specifications remains costly and difficult to automate. Alth

LLM-ADAM: A Generalizable LLM Agent Framework for Pre-Print Anomaly Detection in Additive Manufacturing

Model ReleasesDGX agent

arXiv:2605.03328v1 Announce Type: new Abstract: Additive manufacturing (AM) continues to transform modern manufacturing by enabling flexible, on-demand production of complex geometries across diverse

Low Rank Tensor Completion via Adaptive ADMM

Model ReleasesDGX agent

arXiv:2605.03736v1 Announce Type: cross Abstract: We consider a novel algorithm, for the completion of partially observed low-rank tensors, as a generalization of matrix completion. The proposed low-r

Mantis: Mamba-native Tuning is Efficient for 3D Point Cloud Foundation Models

Model ReleasesDGX agent

arXiv:2605.03438v1 Announce Type: new Abstract: Pre-trained 3D point cloud foundation models (PFMs) have demonstrated strong transferability across diverse downstream tasks. However, full fine-tuning

MAP-Law: Coverage-Driven Retrieval Control for Multi-Turn Legal Consultation

Model ReleasesDGX agent

arXiv:2605.01486v1 Announce Type: new Abstract: Legal consultation is a high-stakes, knowledge-intensive task that requires agents to identify relevant legal issues, retrieve authoritative support, an

Maximizing mutual information between prompts and responses improve LLM personalization with no additional data or human oversight

Model ReleasesDGX agent

arXiv:2603.19294v2 Announce Type: replace-cross Abstract: While post-training has successfully improved large language models (LLMs) across a variety of domains, these gains heavily rely on human-labe

MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following

Model ReleasesDGX agent

arXiv:2605.03858v1 Announce Type: new Abstract: Multi-constraint instruction following requires verifying whether a response satisfies multiple individual requirements, yet LLM judges are often assess

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

Model ReleasesDGX agent

arXiv:2602.00933v2 Announce Type: replace-cross Abstract: The Model Context Protocol (MCP) is rapidly becoming the standard interface for Large Language Models (LLMs) to discover and invoke external t

MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports

Model ReleasesDGX agent

arXiv:2605.03103v1 Announce Type: new Abstract: Semi-structured information extraction (IE) from OCR-derived clinical reports is crucial for efficiently reconstructing patients' longitudinal medical h

Meta-Inverse Physics-Informed Neural Networks for High-Dimensional Ordinary Differential Equations

Model ReleasesDGX agent

arXiv:2605.03511v1 Announce Type: new Abstract: Solving inverse problems in dynamical systems governed by high-dimensional coupled ordinary differential equations (ODEs) is a ubiquitous challenge in s

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models

Model ReleasesDGX agent

arXiv:2605.03485v1 Announce Type: new Abstract: Multidimensional human understanding is essential for real-world applications such as film analysis and virtual digital humans, yet current LVLM benchma

MILE: Mixture of Incremental LoRA Experts for Continual Semantic Segmentation across Domains and Modalities

Model ReleasesDGX agent

arXiv:2605.03555v1 Announce Type: new Abstract: Continual semantic segmentation requires models to adapt to new domains or modalities without sacrificing performance on previously learned tasks. Exper

Mitigating Frequency Learning Bias in Quantum Models via Multi-Stage Residual Learning

Model ReleasesDGX agent

arXiv:2603.10083v2 Announce Type: replace-cross Abstract: Quantum machine learning models based on parameterized circuits can be viewed as Fourier series approximators. However, they often struggle to

Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection

Model ReleasesDGX agent

arXiv:2605.03039v1 Announce Type: new Abstract: Continuous monitoring of bipolar disorder agitation via voice biomarkers requires disentangling stable speaker traits from volatile affective states on

Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability

Model ReleasesDGX agent

arXiv:2605.03217v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model out

Most ReLU Networks Admit Identifiable Parameters

Model ReleasesDGX agent

arXiv:2605.03601v1 Announce Type: new Abstract: We study the realization map of deep ReLU networks, focusing on when a function determines its parameters up to scaling and permutation. To analyze hidd

MRC is already deployed across all of OpenAI’s largest supercomputers that we use to train frontier models, including our site with @Oracle …

Model ReleasesDGX agent

MRC is already deployed across all of OpenAI’s largest supercomputers that we use to train frontier models, including our site with @Oracle Cloud Infrastructure (OCI) in Abilene, Texas, and in @Micros

MSEarth: A Multimodal Benchmark for Earth Science Phenomenon Discovery with MLLMs

Model ReleasesDGX agent

arXiv:2505.20740v3 Announce Type: replace Abstract: The rapid advancement of multimodal large language models (MLLMs) offers new opportunities for complex scientific challenges, yet their application

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling

Model ReleasesDGX agent

arXiv:2605.01566v1 Announce Type: new Abstract: Advances in inference methods have enabled language models to improve their predictions without additional training. These methods often prioritize raw

Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration

Model ReleasesDGX agent

arXiv:2605.03820v1 Announce Type: new Abstract: Multimodal learning often grapples with the challenge of low-quality data, which predominantly manifests as two facets: modality imbalance and noisy cor

Neuro-Symbolic Agents for Hallucination-Free Requirements Reuse

Model ReleasesDGX agent

arXiv:2605.01562v1 Announce Type: cross Abstract: The Object-Oriented Method for Requirements Authoring and Management (OOMRAM) is a requirements reuse framework that relies on exact identifier matchi

NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles

Model ReleasesDGX agent

arXiv:2605.01847v1 Announce Type: new Abstract: Outcome-only evaluation under-specifies whether an evaluated agent profile preserves the commitments required to solve a multi-turn task coherently. Neu

New Bounds for Zarankiewicz Numbers via Reinforced LLM Evolutionary Search

Model ReleasesDGX agent

arXiv:2605.01120v1 Announce Type: new Abstract: The Zarankiewicz number extbf{Z}(m, n, s, t) is the maximum number of edges in a bipartite graph G_{m, n} such that there is no complete K_{s, t} bipart

NEW paper from Microsoft Research. (bookmark it) The entire interpretability literature is built around human readers. As more analysis gets…

Model ReleasesDGX agent

NEW paper from Microsoft Research. (bookmark it) The entire interpretability literature is built around human readers. As more analysis gets delegated to agents, the right target of interpretability s

Not that Groove: Zero-Shot Symbolic Music Editing

Model ReleasesDGX agent

arXiv:2505.08203v2 Announce Type: replace-cross Abstract: While recent advancements in AI music generation have predominantly focused on direct audio synthesis, these systems suffer from inherent rigi

OCRR: A Benchmark for Online Correction Recovery under Distribution Shift

Model ReleasesDGX agent

arXiv:2605.03153v1 Announce Type: cross Abstract: Static benchmarks measure a model frozen at training time. Real systems face distribution shift: new categories, paraphrased queries, drift: and must

omg @bcherny with banger quotes “the future is more async agents… this is why we emphasize verification” “if you’re familiar with higher ord…

Model ReleasesDGX agent

omg @bcherny with banger quotes “the future is more async agents… this is why we emphasize verification” “if you’re familiar with higher order functions, routines are higher order prompts” “default is

On the Spectral Structure and Objective Equivalence of Orthogonal Multilabel Fisher Discriminants

Model ReleasesDGX agent

arXiv:2605.03283v1 Announce Type: cross Abstract: We provide a unified theoretical analysis of Linear Discriminant Analysis with simultaneous multilabel scatter matrix formulations and Stiefel orthogo

On Verbalized Confidence Scores for LLMs

Model ReleasesDGX agent

arXiv:2412.14737v2 Announce Type: replace Abstract: The rise of large language models (LLMs) and their tight integration into our daily life make it essential to dedicate efforts towards their trustwo

OpenAI’s new GPT-5.5 Instant makes ChatGPT smarter, with more concise and reliable responses

Model ReleasesDGX agent

OpenAI Group PBC is replacing the default model in ChatGPT with the launch of GPT-5.5 Instant, claiming users will notice fewer hallucinations when it’s discussing “sensitive topics” such as finance,

Optimal control of the future via prospective learning with control

Model ReleasesDGX agent

arXiv:2511.08717v4 Announce Type: replace-cross Abstract: Optimal control of the future is the next frontier for AI. Current approaches to this problem are typically rooted in reinforcement learning (

ORPilot: A Production-Oriented Agentic LLM-for-OR Tool for Optimization Modeling

Model ReleasesDGX agent

arXiv:2605.02728v1 Announce Type: new Abstract: This paper presents ORPilot, an open-source agentic AI system that translates real-world business problems into solver-ready optimization models. Unlike

Pairwise matrices for sparse autoencoders: single-feature inspection mislabels causal axes

Model ReleasesDGX agent

arXiv:2605.03160v1 Announce Type: new Abstract: The standard sparse-autoencoder (SAE) interpretability protocol labels each feature from its top-activating contexts and validates by single-feature ste

Parameter-Efficient Distributional RL via Normalizing Flows and a Geometry-Aware Cramer Surrogate

Model ReleasesDGX agent

arXiv:2505.04310v2 Announce Type: replace-cross Abstract: Distributional Reinforcement Learning (DistRL) improves upon expectation-based methods by modeling full return distributions, but standard app

Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback

Model ReleasesDGX agent

arXiv:2605.03848v1 Announce Type: new Abstract: Estimating how well a person performs an action, rather than which action is performed, is central to coaching, rehabilitation, and talent identificatio

PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination

Model ReleasesDGX agent

arXiv:2605.03571v1 Announce Type: new Abstract: Patent examination is a complex, multi-stage process requiring both technical expertise and legal reasoning, increasingly challenged by rising applicati

PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs

Model ReleasesDGX agent

arXiv:2605.01123v1 Announce Type: new Abstract: Large language models (LLMs) can provide automated feedback in educational settings, but aligning an LLMs style with a specific instructors tone while m

PHBench: A Benchmark for Predicting Startup Series A Funding from Product Hunt Launch Signals

Model ReleasesDGX agent

arXiv:2605.02974v1 Announce Type: cross Abstract: Structured launch signals on Product Hunt contain statistically significant predictive information for Series A funding outcomes. We construct PHBench

← Previous
1…281282283284285…376
Next →