AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
Human
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “openai”

GridTimelineEvolution
61+ results
7 Jul 2026

How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krugel, and Uhl (2025)

SafetyDGX agent

arXiv:2603.22730v2 Announce Type: replace Abstract: Pfeffer, Krugel, and Uhl (2025) report that OpenAI's reasoning model o1-mini produces more utilitarian responses to the trolley problem and footbrid

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning

Model ReleasesDGX agent

arXiv:2603.23483v2 Announce Type: replace-cross Abstract: Agentic multimodal large language models (MLLMs) (e.g., OpenAI o3 and Gemini Agentic Vision) achieve remarkable reasoning capabilities through

5 Aug 2026

Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2608.02616v1 Announce Type: cross Abstract: We present the first independent, systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, across 42 synth

How Closely Do LLM Reviews Align with Human Peer Review?

Model ReleasesDGX agent

arXiv:2608.03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers

5 May 2026

OpenAI GPT-5 System Card

Model ReleasesDGX agent

arXiv:2601.03267v2 Announce Type: replace Abstract: This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers

1 May 2026

OpenAI o1 System Card

SafetyDGX agent

arXiv:2412.16720v2 Announce Type: replace Abstract: The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provi

In-Context Prompting Obsoletes Agent Orchestration for Procedural Tasks

AgentsDGX agent

arXiv:2604.27891v1 Announce Type: new Abstract: Agent orchestration frameworks -- LangGraph, CrewAI, Google ADK, OpenAI Agents SDK, and others -- place an external orchestrator above the LLM, tracking

HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats

Model ReleasesDGX agent

arXiv:2604.27470v1 Announce Type: new Abstract: Millions of clinicians use ChatGPT to support clinical care, but evaluations of the most common use cases in model-clinician conversations are limited.

28 Jul 2026

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

Model ReleasesDGX agent

arXiv:2509.02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions

26 May 2026

How Well Do Models Follow Their Constitutions?

Model ReleasesDGX agent

arXiv:2605.24229v1 Announce Type: new Abstract: Frontier AI developers now train models against long written behavioral specifications, such as Anthropic's constitution (Anthropic, 2025a) and OpenAI's

MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional

Model ReleasesDGX agent

arXiv:2605.24699v1 Announce Type: new Abstract: Most reported gains on agentic-LLM clinical benchmarks are often attributed to prompt engineering, yet our results suggest that larger improvements can

9 Jul 2026

Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies

Model ReleasesDGX agent

arXiv:2607.06963v1 Announce Type: cross Abstract: Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic

14 Apr 2026

Sanity Checks for Agentic Data Science

AgentsDGX agent

arXiv:2604.11003v1 Announce Type: new Abstract: Agentic data science (ADS) pipelines have grown rapidly in both capability and adoption, with systems such as OpenAI Codex now able to directly analyze

Locket: Robust Feature-Locking Technique for Language Models

ResearchDGX agent

arXiv:2510.12117v3 Announce Type: replace-cross Abstract: Chatbot service providers (e.g., OpenAI) rely on tiered subscription plans to generate revenue, offering black-box access to basic models for

30 Jul 2026

The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness

Model ReleasesDGX agent

arXiv:2503.10647v2 Announce Type: replace Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dim

24 Jul 2026

Autonomous disproofs of the sum-product conjecture over mathbb R with GPT-5.5 Pro

Model ReleasesDGX agent

arXiv:2607.20525v1 Announce Type: new Abstract: OpenAI's recent disproof of the Erdos unit distance conjecture marked a milestone for AI in mathematics. It also inspired another breakthrough: a human

15 Jul 2026

Optimization Is Not All You Need

Model ReleasesDGX agent

arXiv:2607.11977v1 Announce Type: new Abstract: In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text. The alignment that produce

25 Jun 2026

Real-Time Voice AI Hears but Does Not Listen

Model ReleasesDGX agent

arXiv:2606.26083v1 Announce Type: new Abstract: Speech conveys information through both words and vocal delivery. We evaluate four leading production realtime voice systems-OpenAI's GPT Realtime 2, Go

24 Jun 2026

IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPO

Model ReleasesDGX agent

arXiv:2606.23032v2 Announce Type: replace Abstract: Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier language model

23 Jun 2026

AIR: Adaptive Interleaved Reasoning with Code in MLLMs

AgentsDGX agent

arXiv:2606.23678v1 Announce Type: new Abstract: Following the paradigm shift initiated by OpenAI o3, interleaved reasoning with code to enhance multimodal large language models (MLLMs) has become a pi

9 Jun 2026

Subtitle-Aligned Fine-Tuning of Whisper for Swiss German ASR: Benchmark Contamination, Convention Mismatch, and an Honest Baseline at 25.6% WER (13.8% cWER)

Model ReleasesDGX agent

arXiv:2606.07608v1 Announce Type: cross Abstract: We present a systematic study of fine-tuning OpenAI's Whisper large-v3 for Swiss German ASR, using 1,367 hours of broadcast speech paired with Standar

2 Jun 2026

SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning

Model ReleasesDGX agent

arXiv:2603.08000v2 Announce Type: replace Abstract: Large reasoning models (LRMs) like OpenAI o1 and DeepSeek-R1 achieve high accuracy on complex tasks by adopting long chain-of-thought (CoT) reasonin

Dynamic Coordination Strategy Selection for Enterprise Multi-Agent Systems

SafetyDGX agent

arXiv:2606.00804v1 Announce Type: cross Abstract: Enterprise multi-agent systems increasingly expose multiple coordination patterns, but deployments often lack evidence for when to use consensus, deba

29 May 2026

First head-to-head comparison of agentic AI applied to the analysis of simulated data of the Einstein Telescope

Model ReleasesDGX agent

arXiv:2605.28916v1 Announce Type: cross Abstract: We report a comparison of two state-of-the-art agentic AI systems, Claude Code (Anthropic) and Codex (OpenAI), tasked with autonomously executing a si

Persona Conditioning of Brand Recommendations in Retrieval-Augmented Commercial Chat: A Prominence-Stratified Cross-Provider Audit

ApplicationsDGX agent

arXiv:2605.30207v1 Announce Type: new Abstract: The same prompt -- 'best CRM software' -- reaches AI assistants from buyers in widely different contexts: a solo founder, an enterprise VP, a UK SMB own

23 May 2026

Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost

Model ReleasesDGX agent

arXiv:2605.22502v1 Announce Type: cross Abstract: Agent orchestration frameworks have proliferated, collectively exceeding 290,000 GitHub stars across LangGraph, CrewAI, Google ADK, OpenAI Agents SDK,

12 May 2026

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models

Model ReleasesDGX agent

arXiv:2605.09806v1 Announce Type: cross Abstract: Large reasoning models, such as OpenAI o1 and DeepSeek-R1, tend to become increasingly verbose as their reasoning capabilities improve. These inflated

Towards a Large Language-Vision Question Answering Model for MSTAR Automatic Target Recognition

Model ReleasesDGX agent

arXiv:2605.10772v1 Announce Type: cross Abstract: Large language-vision models (LLVM), such as OpenAI's ChatGPT and GPT-4, have gained prominence as powerful tools for analyzing text and imagery. The

7 May 2026

TSCG: Deterministic Tool-Schema Compilation for Agentic LLM Deployments

Model ReleasesDGX agent

arXiv:2605.04107v1 Announce Type: cross Abstract: Production agent frameworks (OpenAI Function Calling, Anthropic Tool Use, MCP) transmit tool schemas as JSON, a format designed for machine parsing, n

Beyond Public Access in LLM Pre-Training Data

SafetyDGX agent

arXiv:2505.00020v2 Announce Type: replace Abstract: Using a legally obtained dataset of 34 copyrighted O'Reilly Media books, we apply the DE-COP membership inference attack method to investigate wheth

30 Apr 2026

When to Retrieve During Reasoning: Adaptive Retrieval for Large Reasoning Models

Model ReleasesDGX agent

arXiv:2604.26649v1 Announce Type: cross Abstract: Large reasoning models such as DeepSeek-R1 and OpenAI o1 generate extended chains of thought spanning thousands of tokens, yet their integration with

Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex

Model ReleasesDGX agent

arXiv:2604.14858v2 Announce Type: replace Abstract: As agent systems move into increasingly diverse execution settings, trajectory-level safety evaluation and diagnosis require benchmarks that evolve

29 Apr 2026

GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment

Model ReleasesDGX agent

arXiv:2604.25370v1 Announce Type: new Abstract: The release of GPT-image-2 by OpenAI marks a watershed moment in AI-generated imagery: the boundary between photographic reality and synthetic content h

When the Forger Is the Judge: GPT-Image-2 Cannot Recognize Its Own Faked Documents

Model ReleasesDGX agent

arXiv:2604.25213v1 Announce Type: new Abstract: OpenAI's GPT-Image-2 has effectively erased the visual boundary between authentic and AI-edited document images: a single number on a receipt can be rep

28 Apr 2026

LongFlow: Efficient KV Cache Compression for Reasoning Models

Model ReleasesDGX agent

arXiv:2603.11504v2 Announce Type: replace-cross Abstract: Recent reasoning models such as OpenAI-o1 and DeepSeek-R1 have shown strong performance on complex tasks including mathematical reasoning and

10 Apr 2026

Clickbait detection: quick inference with maximum impact

ResearchDGX agent

arXiv:2604.08148v1 Announce Type: new Abstract: We propose a lightweight hybrid approach to clickbait detection that combines OpenAI semantic embeddings with six compact heuristic features capturing s

Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts?

Model ReleasesDGX agent

arXiv:2503.18018v2 Announce Type: replace Abstract: We demonstrate that large language models' (LLMs) mathematical reasoning is culturally sensitive: testing 14 models from Anthropic, OpenAI, Google,

More Capable, Less Cooperative? When LLMs Fail At Zero-Cost Collaboration

AgentsDGX agent

arXiv:2604.07821v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly coordinate in multi-agent systems, yet we lack an understanding of where and why cooperation failures m

11 Aug 2026

From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations

Model ReleasesDGX agent

arXiv:2608.07523v1 Announce Type: cross Abstract: Difficulty differences across parallel-class programming examinations affect the fairness of course assessment. This study repositions large language

When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains

Model ReleasesDGX agent

arXiv:2608.07538v1 Announce Type: new Abstract: As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably

An Expectation-Maximization Perspective on Reinforcement Learning for LLM Reasoning

Model ReleasesDGX agent

arXiv:2504.18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated b

AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups

Local AiDGX agent

arXiv:2608.08883v1 Announce Type: new Abstract: Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific workflows. H

Curriculum Generation under Structured Parametric Environments for Robust Navigation Policies

Model ReleasesDGX agent

arXiv:2608.08545v1 Announce Type: cross Abstract: Robust navigation policies for autonomous agents must generalize across continuously varying environmental conditions such as turn rates, obstacles, f

DevIntent: How Much Does LLM-Generated Code Violate Developer Intent?

Model ReleasesDGX agent

arXiv:2608.07614v1 Announce Type: cross Abstract: Code generated by LLMs can violate a developer's implicit intentions when given an ambiguous prompt, yet standard benchmarks measure only whether code

Stealing Reasoning Traces from Proprietary LLM APIs

SafetyDGX agent

arXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim

Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual Reasoning

Local AiDGX agent

arXiv:2608.07593v1 Announce Type: cross Abstract: Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose to eat. Exi

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

Model ReleasesDGX agent

arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha

6 Aug 2026

Document Optimization for Black-Box Retrieval via Reinforcement Learning

ResearchDGX agent

arXiv:2604.05087v3 Announce Type: replace Abstract: Document expansion is a classical technique for improving retrieval quality, and is attractive since it shifts computation offline, avoiding additio

4 Aug 2026

Generative AI and Foundation Models in Medical Image

Model ReleasesDGX agent

arXiv:2608.01686v1 Announce Type: new Abstract: In recent years, generative AI has attracted significant public attention, and its use has been rapidly expanding across a wide range of domains. From c

27 Jul 2026

Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images

Model ReleasesDGX agent

arXiv:2607.21672v1 Announce Type: cross Abstract: Long source-code contexts consume many text tokens, motivating the proposal to render code as images for vision-language models. Recent work asks whet

26 Jun 2026

Dream machine -- the next creative economy

SafetyDGX agent

arXiv:2606.26114v1 Announce Type: cross Abstract: We examine the structural transformation of creative industries under generative artificial intelligence, drawing on 374 primary sources spanning poli

3 Jun 2026

vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models

Model ReleasesDGX agent

arXiv:2603.04444v3 Announce Type: replace-cross Abstract: As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing -- se

27 May 2026

Hubness, Not Anisotropy, Drives Cross-Lingual Retrieval Asymmetry in Multilingual Embedding Models

Model ReleasesDGX agent

arXiv:2605.26575v1 Announce Type: new Abstract: Multilingual embedding models are deployed under the assumption that cross-lingual retrieval is symmetric: if a query in language A retrieves its transl

22 May 2026

AMEL: Accumulated Message Effects on LLM Judgments

Model ReleasesDGX agent

arXiv:2605.22714v1 Announce Type: cross Abstract: Large language models are routinely used as automated evaluators: to review code, moderate content, or score outputs, often with many items passing th

18 May 2026

DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

SafetyDGX agent

arXiv:2511.19399v3 Announce Type: replace-cross Abstract: Deep research agents perform multi-step research to produce long-form, well-attributed answers. However, most open deep research agents are tr

15 Apr 2026

Local-Splitter: A Measurement Study of Seven Tactics for Reducing Cloud LLM Token Usage on Coding-Agent Workloads

Local AiDGX agent

arXiv:2604.12301v1 Announce Type: cross Abstract: We present a systematic measurement study of seven tactics for reducing cloud LLM token usage when a small local model can act as a triage layer in fr

When Reasoning Models Hurt Behavioral Simulation: A Solver-Sampler Mismatch in Multi-Agent LLM Negotiation

Model ReleasesDGX agent

arXiv:2604.11840v1 Announce Type: new Abstract: Large language models are increasingly used as agents in social, economic, and policy simulations. A common assumption is that stronger reasoning should

13 Apr 2026

LLM-Rosetta: A Hub-and-Spoke Intermediate Representation for Cross-Provider LLM API Translation

ApplicationsDGX agent

arXiv:2604.09360v1 Announce Type: cross Abstract: The rapid proliferation of Large Language Model (LLM) providers--each exposing proprietary API formats--has created a fragmented ecosystem where appli

12 Aug 2026

Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

Model ReleasesDGX agent

arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models

10 Aug 2026

Blast Radius

AgentsDGX agent

arXiv:2608.07440v1 Announce Type: new Abstract: Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer that estimates

← Previous
1
Next →
227 results
← Previous
1234
Next →