AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “openai”

GridTimelineEvolution
227 results
Model Releases

When Reasoning Models Hurt Behavioral Simulation: A Solver-Sampler Mismatch in Multi-Agent LLM Negotiation

DGX agent

arXiv:2604.11840v1 Announce Type: new Abstract: Large language models are increasingly used as agents in social, economic, and policy simulations. A common assumption is that stronger reasoning should

model-releasesarxiv-cs-lg
15 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Applications

LLM-Rosetta: A Hub-and-Spoke Intermediate Representation for Cross-Provider LLM API Translation

DGX agent

arXiv:2604.09360v1 Announce Type: cross Abstract: The rapid proliferation of Large Language Model (LLM) providers--each exposing proprietary API formats--has created a fragmented ecosystem where appli

applicationsarxiv-cs-ai
13 Apr 2026
Agents

More Capable, Less Cooperative? When LLMs Fail At Zero-Cost Collaboration

DGX agent

arXiv:2604.07821v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly coordinate in multi-agent systems, yet we lack an understanding of where and why cooperation failures m

agentsarxiv-cs-cl
10 Apr 2026
Model Releases

Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

DGX agent

arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

An Expectation-Maximization Perspective on Reinforcement Learning for LLM Reasoning

DGX agent

arXiv:2504.18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated b

model-releasesarxiv-cs-ai
11 Aug 2026
Local Ai

AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups

DGX agent

arXiv:2608.08883v1 Announce Type: new Abstract: Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific workflows. H

local-aiarxiv-cs-ai
11 Aug 2026
Model Releases

Curriculum Generation under Structured Parametric Environments for Robust Navigation Policies

DGX agent

arXiv:2608.08545v1 Announce Type: cross Abstract: Robust navigation policies for autonomous agents must generalize across continuously varying environmental conditions such as turn rates, obstacles, f

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

DevIntent: How Much Does LLM-Generated Code Violate Developer Intent?

DGX agent

arXiv:2608.07614v1 Announce Type: cross Abstract: Code generated by LLMs can violate a developer's implicit intentions when given an ambiguous prompt, yet standard benchmarks measure only whether code

model-releasesarxiv-cs-cl
11 Aug 2026
Safety

Stealing Reasoning Traces from Proprietary LLM APIs

DGX agent

arXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim

safetyarxiv-cs-ai
11 Aug 2026
Local Ai

Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual Reasoning

DGX agent

arXiv:2608.07593v1 Announce Type: cross Abstract: Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose to eat. Exi

local-aiarxiv-cs-ai
11 Aug 2026
Model Releases

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

DGX agent

arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha

model-releasesarxiv-cs-ai
11 Aug 2026
Agents

Blast Radius

DGX agent

arXiv:2608.07440v1 Announce Type: new Abstract: Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer that estimates

agentsarxiv-cs-ai
10 Aug 2026
Model Releases

Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation

DGX agent

arXiv:2608.06955v1 Announce Type: new Abstract: Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs sys

model-releasesarxiv-cs-ai
10 Aug 2026
Research

Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram Treebanks

DGX agent

arXiv:2608.07283v1 Announce Type: new Abstract: Grammar engineering requires expertise in linguistic formalism and computational implementation, especially in parallel grammar projects that balance cr

researcharxiv-cs-cl
10 Aug 2026
Safety

LMM Modality Transfer: A Pre-requisite for Autonomous GIS Agents

DGX agent

arXiv:2608.06948v1 Announce Type: new Abstract: AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial t

safetyarxiv-cs-ai
10 Aug 2026
Model Releases

NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs

DGX agent

arXiv:2608.07167v1 Announce Type: new Abstract: Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing something it should

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

Social World Models

DGX agent

arXiv:2509.00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about others' perspectives, even with limited informat

model-releasesarxiv-cs-ai
10 Aug 2026
Research

ProDVI: Programmatic Dynamics Priors for Value Network Initialization

DGX agent

arXiv:2608.06015v1 Announce Type: cross Abstract: Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch,

researcharxiv-cs-ai
7 Aug 2026
Model Releases

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

DGX agent

arXiv:2608.06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet m

model-releasesarxiv-cs-ai
7 Aug 2026
Model Releases

Can Post-Training Transform LLMs into Causal Reasoners?

DGX agent

arXiv:2602.06337v2 Announce Type: replace-cross Abstract: Causal inference is essential for decision-making but remains challenging for non-experts. While large language models (LLMs) show promise in

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference

DGX agent

arXiv:2608.02947v1 Announce Type: cross Abstract: The attention score with rotary position embeddings (RoPE) decomposes exactly into a sum over its 2D-rotation frequency pairs, and each pair's wavelen

model-releasesarxiv-cs-cl
5 Aug 2026
Research

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

DGX agent

arXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how m

researcharxiv-cs-cl
5 Aug 2026
Model Releases

Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

DGX agent

arXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive

model-releasesarxiv-cs-cl
5 Aug 2026
Local Ai

SAT-Edge-Agent: Hardware-in-the-Loop Edge-Agent Orchestration for Onboard Satellite Intelligence

DGX agent

arXiv:2608.03728v1 Announce Type: new Abstract: Onboard satellite intelligence requires a task layer that translates mission intent into local tool calls, exposes execution state, and returns machine-

local-aiarxiv-cs-ai
5 Aug 2026
Model Releases

Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

DGX agent

arXiv:2606.15673v2 Announce Type: replace Abstract: Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and of

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

XSPA: Crafting Imperceptible X-Shaped Sparse Adversarial Perturbations for Transferable Attacks on VLMs

DGX agent

arXiv:2603.28568v2 Announce Type: replace Abstract: Vision-language models (VLMs) share visual-textual representations across zero-shot classification, image captioning, and visual question answering

model-releasesarxiv-cs-cv
4 Aug 2026
Local Ai

Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift

DGX agent

arXiv:2607.28696v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) can retain high observed marginal coverage after clinical shift while substantially under-covering an individual

local-aiarxiv-cs-cv
3 Aug 2026
Model Releases

LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

DGX agent

arXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Using Large Language Models for Idea Generation in Innovation

DGX agent

arXiv:2607.27553v1 Announce Type: cross Abstract: This research evaluates the efficacy of large language models (LLMs) in generating new product ideas. To do so, we compare three pools of ideas for ne

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

DGX agent

arXiv:2607.26062v1 Announce Type: cross Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intell

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model

DGX agent

arXiv:2607.22951v1 Announce Type: cross Abstract: Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified ope

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2

DGX agent

arXiv:2606.19259v2 Announce Type: replace-cross Abstract: Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information. As recent multimodal image generation model

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research

DGX agent

arXiv:2607.24032v1 Announce Type: new Abstract: Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages

DGX agent

arXiv:2607.24276v1 Announce Type: cross Abstract: Large language models (LLMs) process text through subword tokenizers rather than directly reading characters or words. Because these tokenizers are tr

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents

DGX agent

arXiv:2607.20468v1 Announce Type: new Abstract: AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or n

model-releasesarxiv-cs-ai
24 Jul 2026
Research

Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit

DGX agent

arXiv:2512.10092v2 Announce Type: replace Abstract: Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors or biases

researcharxiv-cs-ai
24 Jul 2026
Model Releases

IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests

DGX agent

arXiv:2607.20759v1 Announce Type: cross Abstract: AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute code with au

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

DGX agent

arXiv:2607.21518v1 Announce Type: new Abstract: Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction.

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Economic Evaluations of Language Models

DGX agent

arXiv:2607.19375v1 Announce Type: cross Abstract: Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuable task. We

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

When Shippers Become Algorithms: Candidate Exposure, Information Design, and the Concentration of LLM-Mediated Freight Markets

DGX agent

arXiv:2607.19967v1 Announce Type: cross Abstract: Shippers are beginning to delegate carrier selection to large language model (LLM) agents. We ask what such delegation does to a freight matching mark

model-releasesarxiv-cs-ai
23 Jul 2026
Research

When Audio Separation Hurts Zero-Shot ASR: Evaluating SAM-Audio with Whisper on Bengali and English Speech

DGX agent

arXiv:2603.04710v2 Announce Type: replace-cross Abstract: Recent advances in automatic speech recognition (ASR) and speech enhancement have strengthened the common belief that cleaner audio should lea

researcharxiv-cs-ai
16 Jul 2026
Model Releases

DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents

DGX agent

arXiv:2509.21842v2 Announce Type: replace Abstract: Travel planning (TP) agent has recently worked as an emerging building block to interact with external tools/resources for travel itinerary generati

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

So Many Opinions, So Many LLMs: Comparing Large Language Models to Traditional Machine Learning for Open- Ended Survey Analysis

DGX agent

arXiv:2607.11890v1 Announce Type: cross Abstract: Open-ended surveys offer valuable insights, but they are notoriously difficult to analyze at scale. Building on previous work that employed traditiona

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

DGX agent

arXiv:2607.08497v1 Announce Type: cross Abstract: Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, t

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

InfraQR: Edge-Placed QR-Inspired Structured Patch Attacks on Infrared Vision-Language Models

DGX agent

arXiv:2607.07288v1 Announce Type: new Abstract: Infrared vision-language models are increasingly used for perception under low-light and adverse visual conditions, yet their robustness to localized st

model-releasesarxiv-cs-cv
9 Jul 2026
Agents

A Three-Layer Framework for AI in Scientific Discovery

DGX agent

arXiv:2606.13566v2 Announce Type: replace Abstract: Current discussions of AI in scientific discovery are often dominated by two visible capabilities: search over existing knowledge and execution thro

agentsarxiv-cs-ai
8 Jul 2026
Applications

Abductive Corroboration of Probabilistic AI Models for Forensic Synthetic Media Detection

DGX agent

arXiv:2607.05434v1 Announce Type: cross Abstract: Artificial Intelligence (AI) models, at their core, apply general learnings from broad datasets to individual circumstances using probabilistic behavi

applicationsarxiv-cs-cv
8 Jul 2026
Model Releases

From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation

DGX agent

arXiv:2603.15600v2 Announce Type: replace-cross Abstract: Accurate process supervision remains a critical challenge for long-horizon robotic manipulation. A primary bottleneck is that current video ML

model-releasesarxiv-cs-ai
8 Jul 2026
← Previous
12345
Next →