AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “openai”

GridTimelineEvolution
49+ results
Safety

How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krugel, and Uhl (2025)

DGX agent

arXiv:2603.22730v2 Announce Type: replace Abstract: Pfeffer, Krugel, and Uhl (2025) report that OpenAI's reasoning model o1-mini produces more utilitarian responses to the trolley problem and footbrid

safetyarxiv-cs-cl
7 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks

DGX agent

arXiv:2608.02616v1 Announce Type: cross Abstract: We present the first independent, systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, across 42 synth

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

OpenAI GPT-5 System Card

DGX agent

arXiv:2601.03267v2 Announce Type: replace Abstract: This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers

model-releasesarxiv-cs-cl
5 May 2026
Safety

OpenAI o1 System Card

DGX agent

arXiv:2412.16720v2 Announce Type: replace Abstract: The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provi

safetyarxiv-cs-ai
1 May 2026
Model Releases

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

DGX agent

arXiv:2509.02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

How Well Do Models Follow Their Constitutions?

DGX agent

arXiv:2605.24229v1 Announce Type: new Abstract: Frontier AI developers now train models against long written behavioral specifications, such as Anthropic's constitution (Anthropic, 2025a) and OpenAI's

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies

DGX agent

arXiv:2607.06963v1 Announce Type: cross Abstract: Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic

model-releasesarxiv-cs-ai
9 Jul 2026
Agents

Sanity Checks for Agentic Data Science

DGX agent

arXiv:2604.11003v1 Announce Type: new Abstract: Agentic data science (ADS) pipelines have grown rapidly in both capability and adoption, with systems such as OpenAI Codex now able to directly analyze

agentsarxiv-cs-ai
14 Apr 2026
Model Releases

The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness

DGX agent

arXiv:2503.10647v2 Announce Type: replace Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dim

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Autonomous disproofs of the sum-product conjecture over mathbb R with GPT-5.5 Pro

DGX agent

arXiv:2607.20525v1 Announce Type: new Abstract: OpenAI's recent disproof of the Erdos unit distance conjecture marked a milestone for AI in mathematics. It also inspired another breakthrough: a human

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Optimization Is Not All You Need

DGX agent

arXiv:2607.11977v1 Announce Type: new Abstract: In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text. The alignment that produce

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning

DGX agent

arXiv:2603.23483v2 Announce Type: replace-cross Abstract: Agentic multimodal large language models (MLLMs) (e.g., OpenAI o3 and Gemini Agentic Vision) achieve remarkable reasoning capabilities through

model-releasesarxiv-cs-cl
7 Jul 2026
Model Releases

Real-Time Voice AI Hears but Does Not Listen

DGX agent

arXiv:2606.26083v1 Announce Type: new Abstract: Speech conveys information through both words and vocal delivery. We evaluate four leading production realtime voice systems-OpenAI's GPT Realtime 2, Go

model-releasesarxiv-cs-cl
25 Jun 2026
Model Releases

IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPO

DGX agent

arXiv:2606.23032v2 Announce Type: replace Abstract: Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier language model

model-releasesarxiv-cs-ai
24 Jun 2026
Agents

AIR: Adaptive Interleaved Reasoning with Code in MLLMs

DGX agent

arXiv:2606.23678v1 Announce Type: new Abstract: Following the paradigm shift initiated by OpenAI o3, interleaved reasoning with code to enhance multimodal large language models (MLLMs) has become a pi

agentsarxiv-cs-cv
23 Jun 2026
Model Releases

Subtitle-Aligned Fine-Tuning of Whisper for Swiss German ASR: Benchmark Contamination, Convention Mismatch, and an Honest Baseline at 25.6% WER (13.8% cWER)

DGX agent

arXiv:2606.07608v1 Announce Type: cross Abstract: We present a systematic study of fine-tuning OpenAI's Whisper large-v3 for Swiss German ASR, using 1,367 hours of broadcast speech paired with Standar

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning

DGX agent

arXiv:2603.08000v2 Announce Type: replace Abstract: Large reasoning models (LRMs) like OpenAI o1 and DeepSeek-R1 achieve high accuracy on complex tasks by adopting long chain-of-thought (CoT) reasonin

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

First head-to-head comparison of agentic AI applied to the analysis of simulated data of the Einstein Telescope

DGX agent

arXiv:2605.28916v1 Announce Type: cross Abstract: We report a comparison of two state-of-the-art agentic AI systems, Claude Code (Anthropic) and Codex (OpenAI), tasked with autonomously executing a si

model-releasesarxiv-cs-ai
29 May 2026
Applications

Persona Conditioning of Brand Recommendations in Retrieval-Augmented Commercial Chat: A Prominence-Stratified Cross-Provider Audit

DGX agent

arXiv:2605.30207v1 Announce Type: new Abstract: The same prompt -- 'best CRM software' -- reaches AI assistants from buyers in widely different contexts: a solo founder, an enterprise VP, a UK SMB own

applicationsarxiv-cs-ai
29 May 2026
Model Releases

Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost

DGX agent

arXiv:2605.22502v1 Announce Type: cross Abstract: Agent orchestration frameworks have proliferated, collectively exceeding 290,000 GitHub stars across LangGraph, CrewAI, Google ADK, OpenAI Agents SDK,

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models

DGX agent

arXiv:2605.09806v1 Announce Type: cross Abstract: Large reasoning models, such as OpenAI o1 and DeepSeek-R1, tend to become increasingly verbose as their reasoning capabilities improve. These inflated

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Towards a Large Language-Vision Question Answering Model for MSTAR Automatic Target Recognition

DGX agent

arXiv:2605.10772v1 Announce Type: cross Abstract: Large language-vision models (LLVM), such as OpenAI's ChatGPT and GPT-4, have gained prominence as powerful tools for analyzing text and imagery. The

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TSCG: Deterministic Tool-Schema Compilation for Agentic LLM Deployments

DGX agent

arXiv:2605.04107v1 Announce Type: cross Abstract: Production agent frameworks (OpenAI Function Calling, Anthropic Tool Use, MCP) transmit tool schemas as JSON, a format designed for machine parsing, n

model-releasesarxiv-cs-cl
7 May 2026
Agents

In-Context Prompting Obsoletes Agent Orchestration for Procedural Tasks

DGX agent

arXiv:2604.27891v1 Announce Type: new Abstract: Agent orchestration frameworks -- LangGraph, CrewAI, Google ADK, OpenAI Agents SDK, and others -- place an external orchestrator above the LLM, tracking

agentsarxiv-cs-ai
1 May 2026
Model Releases

When to Retrieve During Reasoning: Adaptive Retrieval for Large Reasoning Models

DGX agent

arXiv:2604.26649v1 Announce Type: cross Abstract: Large reasoning models such as DeepSeek-R1 and OpenAI o1 generate extended chains of thought spanning thousands of tokens, yet their integration with

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment

DGX agent

arXiv:2604.25370v1 Announce Type: new Abstract: The release of GPT-image-2 by OpenAI marks a watershed moment in AI-generated imagery: the boundary between photographic reality and synthetic content h

model-releasesarxiv-cs-cv
29 Apr 2026
Model Releases

When the Forger Is the Judge: GPT-Image-2 Cannot Recognize Its Own Faked Documents

DGX agent

arXiv:2604.25213v1 Announce Type: new Abstract: OpenAI's GPT-Image-2 has effectively erased the visual boundary between authentic and AI-edited document images: a single number on a receipt can be rep

model-releasesarxiv-cs-cv
29 Apr 2026
Model Releases

LongFlow: Efficient KV Cache Compression for Reasoning Models

DGX agent

arXiv:2603.11504v2 Announce Type: replace-cross Abstract: Recent reasoning models such as OpenAI-o1 and DeepSeek-R1 have shown strong performance on complex tasks including mathematical reasoning and

model-releasesarxiv-cs-cl
28 Apr 2026
Research

Locket: Robust Feature-Locking Technique for Language Models

DGX agent

arXiv:2510.12117v3 Announce Type: replace-cross Abstract: Chatbot service providers (e.g., OpenAI) rely on tiered subscription plans to generate revenue, offering black-box access to basic models for

researcharxiv-cs-lg
14 Apr 2026
Research

Clickbait detection: quick inference with maximum impact

DGX agent

arXiv:2604.08148v1 Announce Type: new Abstract: We propose a lightweight hybrid approach to clickbait detection that combines OpenAI semantic embeddings with six compact heuristic features capturing s

researcharxiv-cs-cl
10 Apr 2026
Model Releases

Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts?

DGX agent

arXiv:2503.18018v2 Announce Type: replace Abstract: We demonstrate that large language models' (LLMs) mathematical reasoning is culturally sensitive: testing 14 models from Anthropic, OpenAI, Google,

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations

DGX agent

arXiv:2608.07523v1 Announce Type: cross Abstract: Difficulty differences across parallel-class programming examinations affect the fairness of course assessment. This study repositions large language

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains

DGX agent

arXiv:2608.07538v1 Announce Type: new Abstract: As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably

model-releasesarxiv-cs-ai
11 Aug 2026
Research

Document Optimization for Black-Box Retrieval via Reinforcement Learning

DGX agent

arXiv:2604.05087v3 Announce Type: replace Abstract: Document expansion is a classical technique for improving retrieval quality, and is attractive since it shifts computation offline, avoiding additio

researcharxiv-cs-cl
6 Aug 2026
Model Releases

How Closely Do LLM Reviews Align with Human Peer Review?

DGX agent

arXiv:2608.03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Generative AI and Foundation Models in Medical Image

DGX agent

arXiv:2608.01686v1 Announce Type: new Abstract: In recent years, generative AI has attracted significant public attention, and its use has been rapidly expanding across a wide range of domains. From c

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images

DGX agent

arXiv:2607.21672v1 Announce Type: cross Abstract: Long source-code contexts consume many text tokens, motivating the proposal to render code as images for vision-language models. Recent work asks whet

model-releasesarxiv-cs-cv
27 Jul 2026
Safety

Dream machine -- the next creative economy

DGX agent

arXiv:2606.26114v1 Announce Type: cross Abstract: We examine the structural transformation of creative industries under generative artificial intelligence, drawing on 374 primary sources spanning poli

safetyarxiv-cs-ai
26 Jun 2026
Model Releases

vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models

DGX agent

arXiv:2603.04444v3 Announce Type: replace-cross Abstract: As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing -- se

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

Dynamic Coordination Strategy Selection for Enterprise Multi-Agent Systems

DGX agent

arXiv:2606.00804v1 Announce Type: cross Abstract: Enterprise multi-agent systems increasingly expose multiple coordination patterns, but deployments often lack evidence for when to use consensus, deba

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

Hubness, Not Anisotropy, Drives Cross-Lingual Retrieval Asymmetry in Multilingual Embedding Models

DGX agent

arXiv:2605.26575v1 Announce Type: new Abstract: Multilingual embedding models are deployed under the assumption that cross-lingual retrieval is symmetric: if a query in language A retrieves its transl

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional

DGX agent

arXiv:2605.24699v1 Announce Type: new Abstract: Most reported gains on agentic-LLM clinical benchmarks are often attributed to prompt engineering, yet our results suggest that larger improvements can

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AMEL: Accumulated Message Effects on LLM Judgments

DGX agent

arXiv:2605.22714v1 Announce Type: cross Abstract: Large language models are routinely used as automated evaluators: to review code, moderate content, or score outputs, often with many items passing th

model-releasesarxiv-cs-cl
22 May 2026
Safety

DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

DGX agent

arXiv:2511.19399v3 Announce Type: replace-cross Abstract: Deep research agents perform multi-step research to produce long-form, well-attributed answers. However, most open deep research agents are tr

safetyarxiv-cs-ai
18 May 2026
Safety

Beyond Public Access in LLM Pre-Training Data

DGX agent

arXiv:2505.00020v2 Announce Type: replace Abstract: Using a legally obtained dataset of 34 copyrighted O'Reilly Media books, we apply the DE-COP membership inference attack method to investigate wheth

safetyarxiv-cs-cl
7 May 2026
Model Releases

HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats

DGX agent

arXiv:2604.27470v1 Announce Type: new Abstract: Millions of clinicians use ChatGPT to support clinical care, but evaluations of the most common use cases in model-clinician conversations are limited.

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex

DGX agent

arXiv:2604.14858v2 Announce Type: replace Abstract: As agent systems move into increasingly diverse execution settings, trajectory-level safety evaluation and diagnosis require benchmarks that evolve

model-releasesarxiv-cs-ai
30 Apr 2026
Local Ai

Local-Splitter: A Measurement Study of Seven Tactics for Reducing Cloud LLM Token Usage on Coding-Agent Workloads

DGX agent

arXiv:2604.12301v1 Announce Type: cross Abstract: We present a systematic measurement study of seven tactics for reducing cloud LLM token usage when a small local model can act as a triage layer in fr

local-aiarxiv-cs-ai
15 Apr 2026
← Previous
123…5
Next →
227 results