AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,555 results
11 May 2026

Introducing Daybreak: frontier AI for cyber defenders. Daybreak brings together the most capable OpenAI models, Codex, and our security part…

Model ReleasesDGX agent

Introducing Daybreak: frontier AI for cyber defenders. Daybreak brings together the most capable OpenAI models, Codex, and our security partners to accelerate cyber defense and continuously secure sof

Is Your Prompt Poisoning Code? Defect Induction Rates and Security Mitigation Strategies

Model ReleasesDGX agent

arXiv:2510.22944v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have become indispensable for automated code generation, yet the quality and security of their outputs remain a c

Knowing but Not Correcting: Routine Task Requests Suppress Factual Correction in LLMs


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2605.05957v2 Announce Type: replace Abstract: LLMs reliably correct false claims when presented in isolation, yet when the same claims are embedded in task-oriented requests, they often comply r

LARAG: Link-Aware Retrieval Strategy for RAG Systems in Hyperlinked Technical Documentation

Model ReleasesDGX agent

arXiv:2605.07517v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances the factual grounding of Large Language Models by conditioning their outputs on external documents. Howe

Learning Agent Routing From Early Experience

Model ReleasesDGX agent

arXiv:2605.07180v1 Announce Type: new Abstract: LLM agents achieve strong performance on complex reasoning tasks but incur high latency and compute cost. In practice, many queries fall within the capa

Learning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM Agents

Model ReleasesDGX agent

arXiv:2605.06957v1 Announce Type: new Abstract: We present a dynamic policy-learning approach that combines generalized planning and hierarchical task decomposition for LLM-based agents. Our method, H

Learning Material-Aware Hamiltonian Risk Fields for Safe Navigation

Model ReleasesDGX agent

arXiv:2605.07038v1 Announce Type: new Abstract: Risk-aware navigation should be selective: a policy should expose evasive degrees of freedom only when the local scene admits a lower-risk feasible mane

LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation

Model ReleasesDGX agent

arXiv:2605.07640v1 Announce Type: cross Abstract: Remote sensing lithology interpretation is fundamental to geological surveys, mineral exploration, and regional geological mapping. Unlike general lan

LLM-Based Agents for Competitive Landscape Mapping in Drug Asset Due Diligence

Model ReleasesDGX agent

arXiv:2508.16571v4 Announce Type: replace Abstract: In this paper, we describe and benchmark a competitor-discovery component used within an agentic AI system for fast drug asset due diligence. A comp

Local open-weight AI on a laptop has been improving more than twice as fast as Moore's Law! Between May 2024 and May 2026, the most expensiv…

Model ReleasesDGX agent

Local open-weight AI on a laptop has been improving more than twice as fast as Moore's Law! Between May 2024 and May 2026, the most expensive MacBook Pro you could buy stayed at 128 GB of unified memo

Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate

Model ReleasesDGX agent

arXiv:2605.07342v1 Announce Type: cross Abstract: Compile-pass rate is the dominant evaluation signal for LLM code generation, yet for multi-component domain-specific artifacts it can be actively misl

MAS-Algorithm: A Workflow for Solving Algorithmic Programming Problems with a Multi-Agent System

Model ReleasesDGX agent

arXiv:2605.05949v2 Announce Type: replace Abstract: Algorithmic problem solving serves as a rigorous testbed for evaluating structured reasoning in AI coding systems, as it directly reflects a model's

Mask2Cause: Causal Discovery via Adjacency Constrained Causal Attention

Model ReleasesDGX agent

arXiv:2605.07280v1 Announce Type: cross Abstract: Leveraging deep learning for causal discovery in time series remains challenging because existing neural methods predominantly rely on component-wise

Mathematical Reasoning via Intervention-Based Time-Series Causal Discovery Using LLMs as Concept Mastery Simulators

Model ReleasesDGX agent

arXiv:2605.07600v1 Announce Type: cross Abstract: Recent methods for improving LLM mathematical reasoning, whether through MCTS-based test-time search or causal graph-guided knowledge injection, canno

MathlibPR: Pull Request Merge-Readiness Benchmark for Formal Mathematical Libraries

Model ReleasesDGX agent

arXiv:2605.07147v1 Announce Type: cross Abstract: The ecosystem of Lean and Mathlib has become the de facto standard for large language model (LLM) assisted formal reasoning with remarkable successes

MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.07850v1 Announce Type: cross Abstract: With the rise in scale for deep learning models to billions of parameters, the computational cost of fine-tuning remains a significant barrier to depl

MAVEN: Multi-Agent Verification-Elaboration Network with In-Step Epistemic Auditing

Model ReleasesDGX agent

arXiv:2605.07646v1 Announce Type: cross Abstract: While explicit reasoning trajectories enhance model interpretability, existing paradigms often rely on monolithic chains that lack intermediate verifi

McNdroid: A Longitudinal Multimodal Benchmark for Robust Drift Detection in Android Malware

Model ReleasesDGX agent

arXiv:2605.06894v1 Announce Type: cross Abstract: Machine learning (ML) in real-world systems must contend with concept drift, adversarial actors, and a spectrum of potential features with varying cos

Mean-Pooled Cosine Similarity is Not Length-Invariant: Theory and Cross-Domain Evidence for a Length-Invariant Alternative

Model ReleasesDGX agent

arXiv:2605.07345v1 Announce Type: new Abstract: Mean-pooled cosine similarity is the default metric for comparing neural representations across languages, modalities, and tasks. We establish that this

MedAction: Towards Active Multi-turn Clinical Diagnostic LLMs

Model ReleasesDGX agent

arXiv:2605.07305v1 Announce Type: cross Abstract: Most existing LLM diagnoses are evaluated on static, single-turn settings where complete patient information is provided upfront, an oversimplificatio

MedVIGIL: Evaluating Trustworthy Medical VLMs Under Broken Visual Evidence

Model ReleasesDGX agent

arXiv:2605.07919v1 Announce Type: new Abstract: Medical vision--language models (VLMs) are usually evaluated on intact image--question pairs, but trustworthy clinical use requires a stronger property:

Meet the latest Database Center, now with Gemini-powered fleet intelligence

Model ReleasesDGX agent

Managing a modern database fleet is both a scale and cognitive problem. As database estates grow, the effort required to monitor, troubleshoot, and optimize them often outpaces teams’ capacity, who fi

MELD: Multi-Task Equilibrated Learning Detector for AI-Generated Text

Model ReleasesDGX agent

arXiv:2605.06903v1 Announce Type: cross Abstract: Large language models are now embedded in everyday writing workflows, making reliable AI-generated text detection important for academic integrity, co

MicroBi-ConvLSTM: An Ultra-Lightweight Efficient Model for Human Activity Recognition on Resource Constrained Devices

Model ReleasesDGX agent

arXiv:2602.06523v2 Announce Type: replace Abstract: Human Activity Recognition (HAR) on resource constrained wearables requires models that balance accuracy against strict memory and computational bud

MIND: Monge Inception Distance for Generative Models Evaluation

Model ReleasesDGX agent

arXiv:2605.06797v1 Announce Type: new Abstract: We propose the Monge Inception Distance (MIND), a metric for evaluating generative models that addresses key limitations of the widely adopted Frechet I

MiniAppBench: Evaluating the Shift from Text to Interactive HTML Responses in LLM-Powered Assistants

Model ReleasesDGX agent

arXiv:2603.09652v3 Announce Type: replace Abstract: With the rapid advancement of Large Language Models (LLMs) in code generation, human-AI interaction is evolving from static text responses to dynami

MIPIAD: Multilingual Indirect Prompt Injection Attack Defense with Qwen -- TF-IDF Hybrid and Meta-Ensemble Learning

Model ReleasesDGX agent

arXiv:2605.07269v1 Announce Type: new Abstract: Indirect prompt injection remains a persistent weakness in retrieval-augmented and tool-using LLM systems, and the problem becomes harder to characteris

MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference

Model ReleasesDGX agent

arXiv:2605.07363v1 Announce Type: cross Abstract: DeepSeek Sparse Attention (DSA) sets the state of the art for fine-grained inference-time sparse attention by introducing a learned token-wise indexer

Mitigating Cognitive Bias in RLHF by Altering Rationality

Model ReleasesDGX agent

arXiv:2605.06895v1 Announce Type: new Abstract: How can we make models robust to even imperfect human feedback? In reinforcement learning from human feedback (RLHF), human preferences over model outpu

MobileDev-Bench: A Benchmark for Issue Resolution in Mobile Application Development

Model ReleasesDGX agent

arXiv:2603.24946v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown strong performance on automated software engineering tasks, yet existing benchmarks focus primarily on

Model-Driven Policy Optimization in Differentiable Simulators via Stochastic Exploration

Model ReleasesDGX agent

arXiv:2605.07520v1 Announce Type: new Abstract: Differentiable planning enables gradient-based optimization of decision-making problems by leveraging differentiable models of system dynamics. However,

ModelLens: Finding the Best for Your Task from Myriads of Models

Model ReleasesDGX agent

arXiv:2605.07075v1 Announce Type: new Abstract: The open-source model ecosystem now contains hundreds of thousands of pretrained models, yet picking the best model for a new dataset is increasingly in

Modular Lie Algebraic PDE Control of Multibody Flexible Manipulators

Model ReleasesDGX agent

arXiv:2605.06709v1 Announce Type: new Abstract: This paper addresses PDE-based control for flexible multibody robotic systems, presenting a subsystem-based framework for serial manipulators with arbit

More Thinking, More Bias: Length-Driven Position Bias in Reasoning Models

Model ReleasesDGX agent

arXiv:2605.06672v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning and reasoning-tuned models such as DeepSeek-R1 are commonly assumed to reduce shallow heuristic biases by thinking care

Multi-Objective Constraint Inference using Inverse reinforcement learning

Model ReleasesDGX agent

arXiv:2605.06951v1 Announce Type: new Abstract: Constraint inference is widely considered essential to align reinforcement learning agents with safety boundaries and operational guidelines by observin

MultiSoc-4D: A Benchmark for Diagnosing Instruction-Induced Label Collapse in Closed-Set LLM Annotation of Bengali Social Media

Model ReleasesDGX agent

arXiv:2605.06940v1 Announce Type: new Abstract: Annotation automation via Large Language Models (LLMs) is the core approach for scaling NLP datasets; however, LLM behavior with respect to closed-set i

Muon Dynamics as a Spectral Wasserstein Flow

Model ReleasesDGX agent

arXiv:2604.04891v2 Announce Type: replace-cross Abstract: Gradient normalization stabilizes deep-learning optimization, and spectral normalizations are especially natural for matrix-shaped parameter b

My Mac had less available memory than I expected, turned out the 'claude' Claude Code processes on this machine (running in various terminal…

Model ReleasesDGX agent

My Mac had less available memory than I expected, turned out the 'claude' Claude Code processes on this machine (running in various terminal windows) were consuming ~30GB on their own! The largest one

Narrow Secret Loyalty Dodges Black-Box Audits

Model ReleasesDGX agent

arXiv:2605.06846v1 Announce Type: cross Abstract: Recent work identifies secret loyalties as a distinct threat from standard backdoors. A secret loyalty causes a model to covertly advance the interest

NCL-UoR at SemEval-2026 Task 5: Embedding-Based Methods, Fine-Tuning, and LLMs for Word Sense Plausibility Rating

Model ReleasesDGX agent

arXiv:2603.08256v2 Announce Type: replace Abstract: Word sense plausibility rating requires predicting the human-perceived plausibility of a given word sense on a 1-5 scale in the context of short nar

Neural Neural Scaling Laws

Model ReleasesDGX agent

arXiv:2601.19831v2 Announce Type: replace-cross Abstract: Neural scaling laws predict how language model performance improves with increased training inputs. While aggregate metrics like validation lo

Neural Operators as Efficient Function Interpolators

Model ReleasesDGX agent

arXiv:2605.07792v1 Announce Type: cross Abstract: Neural operators (NOs) are designed to learn maps between infinite-dimensional function spaces. We propose a novel reframing of their use. By introduc

New TIL: I figured out how to use my LLM CLI tool in a shebang line, which means you can write executable scripts in English, or hook up mor…

Model ReleasesDGX agent

Simon Willison discovered how to use an LLM command-line interface tool in Unix shebang lines, enabling the creation of executable scripts written in English or natural language. This technique allows

NPMixer: Hierarchical Neighboring Patch Mixing for Time Series Forecasting

Model ReleasesDGX agent

arXiv:2605.07476v1 Announce Type: new Abstract: Multivariate time series forecasting remains a challenge due to the complexity of local temporal dynamics and global dependencies across multiple variab

NS-Net: Decoupling CLIP Semantic Information through NULL-Space for Generalizable AI-Generated Image Detection

Model ReleasesDGX agent

arXiv:2508.01248v4 Announce Type: replace Abstract: The rapid progress of generative models, such as GANs and diffusion models, has facilitated the creation of highly realistic images, raising growing

NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models

Model ReleasesDGX agent

arXiv:2605.07051v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown good performance on various science educational benchmarks, demonstrating their potential for use in science and

OmicsLM: A Multimodal Large Language Model for Multi-Sample Omics Reasoning

Model ReleasesDGX agent

arXiv:2605.06728v1 Announce Type: cross Abstract: Interpreting transcriptomic data is one of the most common analytical tasks in modern biology. Yet most current models either consume expression profi

On the Invariance and Generality of Neural Scaling Laws

Model ReleasesDGX agent

arXiv:2605.07546v1 Announce Type: new Abstract: Neural scaling laws establish a predictable relationship between model performance and data or compute, offering crucial guidance for resource allocatio

One of the most important properties of LLMs that we take for granted is that newer, bigger models are just better at everything. The AI Lab…

Model ReleasesDGX agent

One of the most important properties of LLMs that we take for granted is that newer, bigger models are just better at everything. The AI Labs are pouring effort into economically valuable fields like

OpenAI Campus Network: Student club interest form

Model ReleasesDGX agent

The OpenAI Campus Network is a program that facilitates student engagement with OpenAI's technology and research on college campuses. This interest form allows students to express interest in starting

OpenAI just released its answer to Claude Mythos

Model ReleasesDGX agent

OpenAI is launching Daybreak, an AI initiative focused on detecting and patching vulnerabilities before attackers find them. Daybreak uses the Codex Security AI agent that launched in March to create

OpenAI launches Daybreak, a cybersecurity initiative integrating AI models and Codex Security to help organizations patch vulnerabilities (Alexey Shabanov/TestingCatalog AI News)

Model ReleasesDGX agent

Alexey Shabanov / TestingCatalog AI News: OpenAI launches Daybreak, a cybersecurity initiative integrating AI models and Codex Security to help organizations patch vulnerabilities — OpenAI launches Da

OpenAI launches DeployCo to help businesses build around intelligence

Model ReleasesDGX agent

OpenAI launched DeployCo, a new service designed to assist businesses in building and deploying applications leveraging OpenAI's AI models and intelligence capabilities. The offering appears to focus

OpenAI launches professional services business with $4B investment

Model ReleasesDGX agent

OpenAI Group PBC today unveiled a new business unit, The OpenAI Deployment Company, that will help companies adopt its artificial intelligence models. The subsidiary is launching with 4 billion in fun

OpenAI launches the OpenAI Deployment Company with a $4B+ investment to help organizations build and deploy AI systems, and acquires AI consulting firm Tomoro (Reuters)

Model ReleasesDGX agent

Reuters: OpenAI launches the OpenAI Deployment Company with a 4B+ investment to help organizations build and deploy AI systems, and acquires AI consulting firm Tomoro — OpenAI said on Monday it is set

Optimal Experiments for Partial Causal Effect Identification

Model ReleasesDGX agent

arXiv:2605.06993v1 Announce Type: new Abstract: Causal queries are often only partially identifiable from observational data, and experiments that could tighten the resulting bounds are typically cost

Optimizing Language Models for Crosslingual Knowledge Consistency

Model ReleasesDGX agent

arXiv:2603.04678v2 Announce Type: replace-cross Abstract: Large language models are known to often exhibit inconsistent knowledge. This is particularly problematic in multilingual scenarios, where mod

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling

Model ReleasesDGX agent

arXiv:2605.07815v1 Announce Type: cross Abstract: Muon improves neural-network training by orthogonalizing matrix-valued updates, but it leaves each layer's update magnitude controlled mostly by a glo

Outlier Smoothing with Closed-Form Rotations for W4A4 Large Language Model Quantization

Model ReleasesDGX agent

arXiv:2511.22316v2 Announce Type: replace Abstract: Large Language Models (LLMs) quantization facilitates deploying LLMs in resource-limited settings, but existing methods that combine incompatible gr

PAIR-Former: Budgeted Relational Multi-Instance Learning for Functional miRNA Target Prediction

Model ReleasesDGX agent

arXiv:2602.00465v3 Announce Type: replace-cross Abstract: Functional miRNA--mRNA targeting is a large-bag prediction problem where each transcript yields a heavy-tailed pool of candidate target sites

← Previous
1…272273274275276…376
Next →