AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,340 results
29 Jul 2026

Install the open-source Codex Security CLI,: npm install @OpenAI/codex-security Or start with: npx @OpenAI/codex-security@latest --help NPM:…

Model ReleasesDGX agent

OpenAI has released the open‑source Codex Security CLI, which can be installed with `npm install @OpenAI/codex-security` or run directly via `npx @OpenAI/codex-security@latest --help`. The tool scans

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications

Model ReleasesDGX agent

arXiv:2607.25642v1 Announce Type: cross Abstract: Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models

Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2607.25292v1 Announce Type: new Abstract: Silicon sampling uses language models as proxies for human survey respondents, treating each model call as an independent draw from the persona's respon

Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do

Model ReleasesDGX agent

arXiv:2607.26015v1 Announce Type: new Abstract: Syntactic convergence (the tendency of speakers to adapt in language towards the grammatical profiles of their interlocutors) is a well-documented featu

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

Model ReleasesDGX agent

arXiv:2607.25904v1 Announce Type: new Abstract: Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluat

JobMatchAI-An Intelligent Job Matching Platform Using Knowledge Graphs, Semantic Search and Explainable AI

Model ReleasesDGX agent

arXiv:2603.14558v3 Announce Type: replace Abstract: Recruiters and job seekers rely on search systems to navigate labor markets, making candidate matching engines critical for hiring outcomes. Most sy

Joint Text-Audio Alignment for EEG-to-Text Decoding in Chinese Speech Production and Perception

Model ReleasesDGX agent

arXiv:2607.25626v1 Announce Type: new Abstract: Decoding speech information directly from scalp electroencephalography (EEG) into text provides a potential non-invasive neural communication pathway fo

Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels

Model ReleasesDGX agent

arXiv:2607.24762v1 Announce Type: new Abstract: Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as mat

LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection

Model ReleasesDGX agent

arXiv:2607.25962v1 Announce Type: new Abstract: Recent generative models can produce images with few obvious visual artifacts, weakening detectors and explanations that rely only on surface appearance

Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks

Model ReleasesDGX agent

arXiv:2607.25257v1 Announce Type: cross Abstract: Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a model's late

Learned, Relied Upon, or Necessary? Separating Checkpoint Dependence from Task-Level Value in Sheaf GNNs

Model ReleasesDGX agent

arXiv:2607.25387v1 Announce Type: new Abstract: Learned restriction maps in sheaf graph neural networks are often treated as proof that the model has discovered useful edge geometry. That conclusion d

Learning from 53.6K Real-World Developer Edits of AI-Generated Code

Model ReleasesDGX agent

arXiv:2607.25130v1 Announce Type: cross Abstract: Imperfections in AI-generated code require that software developers modify the generated code manually, or by re-prompting an AI programming assistant

Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization

Model ReleasesDGX agent

arXiv:2607.25136v1 Announce Type: new Abstract: Research on preference optimization often varies the training objective while holding the data fixed. We instead ask whether a small, high-confidence se

Localized Adaptation Reveals Distinct Learning Signatures in Transformers

Model ReleasesDGX agent

arXiv:2607.25663v1 Announce Type: new Abstract: Transformer adaptation is typically distributed across model depth, even when the intended change is narrow. We investigate how adaptation site shapes w

Localized Anomaly Detection via Differentiable D-vine Copulas

Model ReleasesDGX agent

arXiv:2607.25020v1 Announce Type: new Abstract: Vine copulas provide a flexible framework for modeling complex multivariate distributions through a hierarchical decomposition into bivariate pair-copul

M^2PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation

Model ReleasesDGX agent

arXiv:2510.13434v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with human preferences is pivotal for Machine Translation (MT), yet current approaches are often hindered by m

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Model ReleasesDGX agent

arXiv:2607.24904v1 Announce Type: cross Abstract: Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streamin

Mechanisms of Width Scaling in Normalized Residual Networks: The Effective Alignment Dimension

Model ReleasesDGX agent

arXiv:2607.24887v1 Announce Type: cross Abstract: Existing theories of neural-network width characterize asymptotic limits, but provide limited guidance on whether an expansion direction identified fr

Med-R^3: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning

Model ReleasesDGX agent

arXiv:2507.23541v5 Announce Type: replace Abstract: In medical scenarios, effectively retrieving external knowledge and leveraging it for rigorous logical reasoning is of significant importance. Despi

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing

Model ReleasesDGX agent

arXiv:2607.25300v1 Announce Type: new Abstract: Video editing is fundamentally message-driven: even from the same source footage, the selected shots change depending on the narrative the editor wishes

MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA

Model ReleasesDGX agent

arXiv:2607.24838v1 Announce Type: cross Abstract: In medical multiple-choice question answering (MCQA), Retrieval-Augmented Generation (RAG) can supplement the domain knowledge of language models (LMs

Memory for Large Language Models

Model ReleasesDGX agent

arXiv:2607.25380v1 Announce Type: new Abstract: Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an implicit byproduct of computation to a

MemSFT: Mitigating Alignment Tax with an External Parametric Memory

Model ReleasesDGX agent

arXiv:2607.25614v1 Announce Type: cross Abstract: Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastro

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.25891v1 Announce Type: new Abstract: Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts focus on nar

Microsoft confirms Copilot ‘super app’ coming this year

Model ReleasesDGX agent

Microsoft is working on an AI 'super app' that combines Copilot's chat, coding, and agentic capabilities. During an earnings call on Wednesday, Microsoft CEO Satya Nadella said the app will span 'both

Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models

Model ReleasesDGX agent

arXiv:2607.25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a fluent promp

Model + harness. We have barely begun to understand the best ways to do harness engineering. A huge amount of untapped potential even withou…

Model ReleasesDGX agent

Model + harness. We have barely begun to understand the best ways to do harness engineering. A huge amount of untapped potential even without models getting better (but models are getting better) Turn

MorphUNet: Alpha-Controlled Biometric Transport for Diffusion-Based Face Morphing Attacks

Model ReleasesDGX agent

arXiv:2607.25092v1 Announce Type: new Abstract: Face morphing attacks create synthetic images verifiable against multiple identities, threatening border control and identity verification systems. We i

Multi-Fidelity Learning with Shallow Recurrent Decoders for Multi-Physics Applications

Model ReleasesDGX agent

arXiv:2606.05202v2 Announce Type: replace-cross Abstract: In reactor physics, neutronics and multi-physics phenomena can be modelled at different fidelity levels. High-fidelity models based on the Bol

Multi-Scale Structural Features for Continual, Comprehensible Visual Recognition in a Developmental Learning Framework

Model ReleasesDGX agent

arXiv:2607.25531v1 Announce Type: cross Abstract: Contemporary machine learning struggles to learn continually, reuse prior knowledge, and expose a comprehensible internal structure. A recently propos

Multimodal Hybrid Retrieval-Augmented Generation for Scientific Document Understanding using Open-Source SLMs

Model ReleasesDGX agent

arXiv:2607.24799v1 Announce Type: cross Abstract: Large Language Models tend to hallucinate when answering domain-specific ques tions from scientific documents without prior fine-tuning. Currently, me

MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

Model ReleasesDGX agent

arXiv:2607.25186v1 Announce Type: new Abstract: Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, mu

Neurai-VN Benchmark: Standardized Machine Learning Models for Multimodal Digital Phenotyping in Mental Health Classification

Model ReleasesDGX agent

arXiv:2607.25232v1 Announce Type: new Abstract: Digital phenotyping (DP) using smartphones and wearable devices has shown considerable potential for mental health monitoring. However, progress remains

Neuromorphic Diffusion Language Models: Addressing Compute and Memory Bottlenecks via Sparsity and Block Denoising

Model ReleasesDGX agent

arXiv:2607.24841v1 Announce Type: new Abstract: Autoregressive (AR) large language models (LLMs) are inherently inefficient at inference time because each generated token requires accessing the full s

ObliCity: A Benchmark and Baseline for Roof-to-Ground Projection Displacement Correction

Model ReleasesDGX agent

arXiv:2607.25210v1 Announce Type: new Abstract: Oblique-view urban remote sensing imagery inevitably exhibits geometric projection displacements between building roofs and footprints, leading to signi

OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2607.25641v1 Announce Type: cross Abstract: While text-to-image models exhibit remarkable visual fidelity, they frequently violate fundamental physical commonsense. Existing benchmarks often rel

On benchmarking long-context agentic instruction following. Agent benchmarks mostly reward reaching the answer. This new benchmark measures …

Model ReleasesDGX agent

On benchmarking long-context agentic instruction following. Agent benchmarks mostly reward reaching the answer. This new benchmark measures whether the agent reached it the permitted way, which is the

On the Use of LLMs for Specialised Terminology: A Good Alternative to Corpora?

Model ReleasesDGX agent

arXiv:2607.24784v1 Announce Type: new Abstract: Specialised translation relies on the use of documentary and terminological resources, including corpora. These resources are particularly useful for te

OpenPVMapper: A Multi-source, Nationwide Database of Rooftop Photovoltaic Systems in France

Model ReleasesDGX agent

arXiv:2607.25153v1 Announce Type: new Abstract: Rooftop photovoltaic (PV) systems account for the vast majority of PV grid connections, yet no open, comprehensive, installation-level dataset of these

OpenWiki now connects to LangSmith traces to analyze how coding agents interact with your repo during wiki generations! We added a LangSmith…

Model ReleasesDGX agent

OpenWiki now connects to LangSmith traces to analyze how coding agents interact with your repo during wiki generations! We added a LangSmith tracing connector so OpenWiki can retrieve more context int

OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation

Model ReleasesDGX agent

arXiv:2607.25656v1 Announce Type: new Abstract: Complex tasks often decompose into parallelizable yet interdependent subtasks, making orchestration critical to the performance of multi-agent systems (

OrthKD: Extracting Generalized Clinical Knowledge from Heterogeneous Teachers for Lightweight Deployment

Model ReleasesDGX agent

arXiv:2607.25545v1 Announce Type: cross Abstract: Deploying diabetic retinopathy (DR) screening models in primary care requires edge-efficient systems that remain accurate, safe, and reliable under do

PanoLess: Environment Reconstruction from Partial Reflective Views

Model ReleasesDGX agent

arXiv:2607.25362v1 Announce Type: new Abstract: Reflections from shiny objects and glass facades naturally extend the field of view of a camera, capturing the surrounding environment without the need

Parallel Decoding Distillation for Fast Image and Video Generation

Model ReleasesDGX agent

arXiv:2607.26004v1 Announce Type: new Abstract: Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA

Participants will receive access to our frontier models, including our GPT-5.6 family of models. Each workspace includes business-grade priv…

Model ReleasesDGX agent

Participants will receive access to our frontier models, including our GPT-5.6 family of models. Each workspace includes business-grade privacy and security protections. Researcher data is not used to

Pass the Baton: Trajectory-Relayed On-Policy Distillation

Model ReleasesDGX agent

arXiv:2607.26057v1 Announce Type: cross Abstract: On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commit

PATHFinder Agent for Tailored Prenatal Care

Model ReleasesDGX agent

arXiv:2607.24768v1 Announce Type: new Abstract: Prenatal care is an important preventive service designed to improve outcomes for pregnant individuals. The American College of Obstetricians and Gyneco

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

Model ReleasesDGX agent

arXiv:2607.25485v1 Announce Type: new Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Pri

PEANUT: Perturbations by Eigenvector Alignment for Attacking Graph Neural Networks Under Topology-Driven Message Passing

Model ReleasesDGX agent

arXiv:2603.26136v3 Announce Type: replace Abstract: Message Passing Neural Networks (MPNNs) have achieved strong performance on tasks involving relational data. However, small perturbations to graph s

Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

Model ReleasesDGX agent

arXiv:2607.25915v1 Announce Type: new Abstract: Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or b

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2607.24957v1 Announce Type: new Abstract: We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Model

Personalization, Personas, and Forecasting in Value Alignment

Model ReleasesDGX agent

arXiv:2607.24782v1 Announce Type: new Abstract: LLM behavior may be conditioned by human identity in several ways: they may be asked to adapt to users, role-play populations, or forecast how people wo

Physics-Aware End-to-End Deep Reinforcement Learning for Quadcopter Control with Actuator Dynamics

Model ReleasesDGX agent

arXiv:2607.25985v1 Announce Type: cross Abstract: Unmanned aerial vehicles (UAVs), particularly quadcopters, present unique challenges for autonomous control due to their underactuated dynamics: only

Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision

Model ReleasesDGX agent

arXiv:2607.25321v1 Announce Type: new Abstract: Video diffusion models generate visually compelling content but routinely violate elementary physics when the subject involves fluids: liquid columns br

PIcsC: Partitioning-Induced Covariate Shift Correction

Model ReleasesDGX agent

arXiv:2607.25441v1 Announce Type: new Abstract: Covariate shift across training-data partitions biases model selection and parameter estimation in cross-validation, lifelong learning, and federated le

PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning

Model ReleasesDGX agent

arXiv:2508.00344v5 Announce Type: replace Abstract: Large Language Models (LLMs) have shown remarkable advancements in tackling agent-oriented tasks. Despite their potential, existing work faces chall

Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections

Model ReleasesDGX agent

arXiv:2607.25953v1 Announce Type: new Abstract: As LLMs increasingly mediate the political information citizens rely on, there is still no standardized way to assess whether they do so responsibly. We

Probabilistic Symbolic Regression for Equation Discovery via Operator-induced and Regularized Symbolic Forests

Model ReleasesDGX agent

arXiv:2509.19710v2 Announce Type: cross Abstract: Symbolic regression has emerged as a powerful tool for artificial intelligence-driven scientific discovery by learning interpretable analytical expres

PSA: llama.cpp now loads MTP tensors by default for any draft-mtp arch, even with MTP disabled

Model ReleasesDGX agent

If your GGUF has MTP/NextN tensors baked in (GLM-5.2, hy_v3, qwen35moe, step35, etc.), recent llama.cpp builds load them by default — even if you never pass --spec-type draft-mtp. Before, they were sk

Quantizing Kimi K3 (2.8T A50B) to GGUF ourselves - Q3_K_S works, 1.1 TB on disk

Model ReleasesDGX agent

we're experimenting with our own dynamic GGUF quants of kimi k3, made from the original weights with our llama.cpp fork. Q3_K_S is done and works 1114.76 GiB on disk. Q1 and Q2 are in progress, result

← Previous
1…5455565758…373
Next →