AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “openai”

GridTimelineEvolution
227 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Companies

CompanyAnthropic8 recent entries
30 Jul 2026Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

arXiv:2607.26062v1 Announce Type: cross Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intell

→31 Jul 2026LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

arXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or


Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
→5 Aug 2026Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

arXiv:2606.15673v2 Announce Type: replace Abstract: Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and of

→5 Aug 2026How Closely Do LLM Reviews Align with Human Peer Review?

arXiv:2608.03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers

→5 Aug 2026Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

arXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how m

→10 Aug 2026Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation

arXiv:2608.06955v1 Announce Type: new Abstract: Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs sys

→11 Aug 2026Stealing Reasoning Traces from Proprietary LLM APIs

arXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim

→11 Aug 2026DevIntent: How Much Does LLM-Generated Code Violate Developer Intent?

arXiv:2608.07614v1 Announce Type: cross Abstract: Code generated by LLMs can violate a developer's implicit intentions when given an ambiguous prompt, yet standard benchmarks measure only whether code

CompanyOpenAI8 recent entries
11 Aug 2026Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual Reasoning

arXiv:2608.07593v1 Announce Type: cross Abstract: Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose to eat. Exi

→11 Aug 2026Stealing Reasoning Traces from Proprietary LLM APIs

arXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim

→11 Aug 2026From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations

arXiv:2608.07523v1 Announce Type: cross Abstract: Difficulty differences across parallel-class programming examinations affect the fairness of course assessment. This study repositions large language

→11 Aug 2026DevIntent: How Much Does LLM-Generated Code Violate Developer Intent?

arXiv:2608.07614v1 Announce Type: cross Abstract: Code generated by LLMs can violate a developer's implicit intentions when given an ambiguous prompt, yet standard benchmarks measure only whether code

→11 Aug 2026Curriculum Generation under Structured Parametric Environments for Robust Navigation Policies

arXiv:2608.08545v1 Announce Type: cross Abstract: Robust navigation policies for autonomous agents must generalize across continuously varying environmental conditions such as turn rates, obstacles, f

→11 Aug 2026AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups

arXiv:2608.08883v1 Announce Type: new Abstract: Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific workflows. H

→11 Aug 2026An Expectation-Maximization Perspective on Reinforcement Learning for LLM Reasoning

arXiv:2504.18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated b

→12 Aug 2026Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models

CompanyGoogle8 recent entries
27 Jul 2026Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images

arXiv:2607.21672v1 Announce Type: cross Abstract: Long source-code contexts consume many text tokens, motivating the proposal to render code as images for vision-language models. Recent work asks whet

→28 Jul 2026OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

arXiv:2509.02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions

→30 Jul 2026The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness

arXiv:2503.10647v2 Announce Type: replace Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dim

→31 Jul 2026LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

arXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or

→4 Aug 2026Generative AI and Foundation Models in Medical Image

arXiv:2608.01686v1 Announce Type: new Abstract: In recent years, generative AI has attracted significant public attention, and its use has been rapidly expanding across a wide range of domains. From c

→5 Aug 2026How Closely Do LLM Reviews Align with Human Peer Review?

arXiv:2608.03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers

→5 Aug 2026Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

arXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive

→7 Aug 2026ProDVI: Programmatic Dynamics Priors for Value Network Initialization

arXiv:2608.06015v1 Announce Type: cross Abstract: Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch,

CompanyMeta8 recent entries
7 Jul 2026Evaluating Large Language Models for Antisemitic Incident Classification

arXiv:2607.04890v1 Announce Type: new Abstract: Addressing hate and violence in society requires timely detection of hateful events from public reporting, but automated identification of hateful event

→9 Jul 2026Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies

arXiv:2607.06963v1 Announce Type: cross Abstract: Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic

→15 Jul 2026So Many Opinions, So Many LLMs: Comparing Large Language Models to Traditional Machine Learning for Open- Ended Survey Analysis

arXiv:2607.11890v1 Announce Type: cross Abstract: Open-ended surveys offer valuable insights, but they are notoriously difficult to analyze at scale. Building on previous work that employed traditiona

→30 Jul 2026Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

arXiv:2607.26062v1 Announce Type: cross Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intell

→5 Aug 2026ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference

arXiv:2608.02947v1 Announce Type: cross Abstract: The attention score with rotary position embeddings (RoPE) decomposes exactly into a sum over its 2D-rotation frequency pairs, and each pair's wavelen

→7 Aug 2026ProDVI: Programmatic Dynamics Priors for Value Network Initialization

arXiv:2608.06015v1 Announce Type: cross Abstract: Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch,

→10 Aug 2026NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs

arXiv:2608.07167v1 Announce Type: new Abstract: Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing something it should

→11 Aug 2026Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha

CompanyMistral8 recent entries
20 Apr 2026AISysRev -- LLM-based Tool for Title-abstract Screening

arXiv:2510.06708v3 Announce Type: replace-cross Abstract: Conducting systematic reviews is laborious. In the screening or study selection phase, the number of papers can be overwhelming. Recent resear

→12 May 2026Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models

arXiv:2510.08592v3 Announce Type: replace-cross Abstract: Test-Time Scaling (TTS) improves LLM reasoning by exploring multiple candidate responses and then operating over this set to find the best out

→27 May 2026Hubness, Not Anisotropy, Drives Cross-Lingual Retrieval Asymmetry in Multilingual Embedding Models

arXiv:2605.26575v1 Announce Type: new Abstract: Multilingual embedding models are deployed under the assumption that cross-lingual retrieval is symmetric: if a query in language A retrieves its transl

→29 May 2026BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base

arXiv:2605.29379v1 Announce Type: new Abstract: We present BrahmicTokenizer-131K, a 131,072-vocabulary byte-level BPE tokenizer that closes the Brahmic compression gap at the 131K-vocabulary class whi

→24 Jun 2026Transformer-Based Language Models Across Domain Verticals: Architectures, Applications and Critical Assessment

arXiv:2606.24331v1 Announce Type: new Abstract: Transformer-based language models have become the default substrate for natural language processing and the pace of new releases has made it hard for pr

→30 Jul 2026Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

arXiv:2607.26062v1 Announce Type: cross Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intell

→10 Aug 2026Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation

arXiv:2608.06955v1 Announce Type: new Abstract: Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs sys

→11 Aug 2026Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha

CompanyxAI6 recent entries
14 Apr 2026Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness

arXiv:2603.22816v3 Announce Type: replace-cross Abstract: Language models increasingly show their work by writing step-by-step reasoning before answering. But are these steps genuinely used, or is the

→15 Apr 2026Narrative over Numbers: The Identifiable Victim Effect and its Amplification Under Alignment and Reasoning in Large Language Models

arXiv:2604.12076v1 Announce Type: cross Abstract: The Identifiable Victim Effect (IVE) - the tendency to allocate greater resources to a specific, narratively described victim than to a statistically

→23 May 2026A Mechanistic Explanatory Strategy for XAI

arXiv:2411.01332v5 Announce Type: replace Abstract: Despite significant advancements in XAI, scholars note a persistent lack of solid conceptual foundations and integration with broader scientific dis

→9 Jul 2026Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies

arXiv:2607.06963v1 Announce Type: cross Abstract: Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic

→24 Jul 2026Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit

arXiv:2512.10092v2 Announce Type: replace Abstract: Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors or biases

→28 Jul 2026OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

arXiv:2509.02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions

CompanyDeepSeek8 recent entries
24 Jun 2026Transformer-Based Language Models Across Domain Verticals: Architectures, Applications and Critical Assessment

arXiv:2606.24331v1 Announce Type: new Abstract: Transformer-based language models have become the default substrate for natural language processing and the pace of new releases has made it hard for pr

→26 Jun 2026From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

arXiv:2606.26196v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially fo

→30 Jun 2026Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates

arXiv:2606.30085v1 Announce Type: new Abstract: Large-language models have proven to be remarkable if inconsistent parrots of public attitudes and opinions. The extent to which LLMs are able to produc

→15 Jul 2026DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents

arXiv:2509.21842v2 Announce Type: replace Abstract: Travel planning (TP) agent has recently worked as an emerging building block to interact with external tools/resources for travel itinerary generati

→5 Aug 2026Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

arXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive

→11 Aug 2026Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha

→11 Aug 2026An Expectation-Maximization Perspective on Reinforcement Learning for LLM Reasoning

arXiv:2504.18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated b

→12 Aug 2026Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models

CompanyNVIDIA2 recent entries
9 Jun 2026Subtitle-Aligned Fine-Tuning of Whisper for Swiss German ASR: Benchmark Contamination, Convention Mismatch, and an Honest Baseline at 25.6% WER (13.8% cWER)

arXiv:2606.07608v1 Announce Type: cross Abstract: We present a systematic study of fine-tuning OpenAI's Whisper large-v3 for Swiss German ASR, using 1,367 hours of broadcast speech paired with Standar

→11 Aug 2026Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha