AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,512 results
30 Apr 2026

Introducing Agent Collabs: Bring your own ml-interns and agents for collaborative autoresearch! We built a simple platform for swarms of age…

Model ReleasesDGX agent

Introducing Agent Collabs: Bring your own ml-interns and agents for collaborative autoresearch! We built a simple platform for swarms of agents to work together on a problem: they can exchange message

Join us Tue 5/5: #DeepSeek-V4's hybrid attention + sparse MoE reduces KV cache up to 90%, enabling 1M-token context. We'll cover why that ma…

Model ReleasesDGX agent

Join us Tue 5/5: #DeepSeek-V4's hybrid attention + sparse MoE reduces KV cache up to 90%, enabling 1M-token context. We'll cover why that makes it great for agentic workflows, what it took to serve at

L2RU: a Structured State Space Model with prescribed L2-bound


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2503.23818v3 Announce Type: replace-cross Abstract: Structured state-space models (SSMs) have recently emerged as a powerful architecture at the intersection of machine learning and control, fea

LATTICE: Evaluating Decision Support Utility of Crypto Agents

Model ReleasesDGX agent

arXiv:2604.26235v1 Announce Type: cross Abstract: We introduce LATTICE, a benchmark for evaluating the decision support utility of crypto agents in realistic user-facing scenarios. Prior crypto agent

Learning Neural Operator Surrogates for the Black Hole Accretion Code

Model ReleasesDGX agent

arXiv:2604.25985v1 Announce Type: cross Abstract: General-relativistic magnetohydrodynamic (GR-MHD) simulations are essential for studying black hole accretion, relativistic jets, and magnetic reconne

Learning Over-Relaxation Policies for ADMM with Convergence Guarantees

Model ReleasesDGX agent

arXiv:2604.26932v1 Announce Type: cross Abstract: The Alternating Direction Method of Multipliers (ADMM) is a widely used method for structured convex optimization, and its practical performance depen

Learning to Ask: When LLM Agents Meet Unclear Instruction

Model ReleasesDGX agent

arXiv:2409.00557v4 Announce Type: replace-cross Abstract: Equipped with the capability to call functions, modern large language models (LLMs) can leverage external tools for addressing a range of task

lisan say more mean things about us you're being too nice

Model ReleasesDGX agent

lisan say more mean things about us you're being too nice GPT-5.5 is on par with Claude Mythos - GPT-5.5 average pass rate of 71.4% (±8.0%) - Mythos Preview 68.6% (±8.7%) - GPT-5.5 solved a task that

LIT-RAGBench: Benchmarking Generator Capabilities of Large Language Models in Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2603.06198v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) is a framework in which a Generator, such as a Large Language Model (LLM), produces answers by retrieving docum

LLM-Flax : Generalizable Robotic Task Planning via Neuro-Symbolic Approaches with Large Language Models

Model ReleasesDGX agent

arXiv:2604.26569v1 Announce Type: new Abstract: Deploying a neuro-symbolic task planner on a new domain today requires significant manual effort: a domain expert must author relaxation and complementa

LLM Psychosis: A Theoretical and Diagnostic Framework for Reality-Boundary Failures in Large Language Models

Model ReleasesDGX agent

arXiv:2604.25934v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) as interactive agents has exposed a category of behavioral failure that prevailing terminology, princip

Lunguage: A Benchmark for Structured and Sequential Chest X-ray Interpretation

Model ReleasesDGX agent

arXiv:2505.21190v2 Announce Type: replace-cross Abstract: Radiology reports convey detailed clinical observations and capture diagnostic reasoning that evolves over time. However, existing evaluation

LWiAI Podcast #242 - ChatGPT Images 2.0, Qwen 3.6 Max, Kimi-K2.6

Model ReleasesDGX agent

This podcast episode covers recent AI developments including ChatGPT's updated image generation capabilities (Images 2.0), updates to Alibaba's Qwen model reaching version 3.6 Max, and improvements to

Lyapunov-Guided Self-Alignment: Test-Time Adaptation for Offline Safe Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.26516v1 Announce Type: cross Abstract: Offline reinforcement learning (RL) agents often fail when deployed, as the gap between training datasets and real environments leads to unsafe behavi

MARVIS: Modality Adaptive Reasoning over VISualizations

Model ReleasesDGX agent

arXiv:2507.01544v2 Announce Type: replace Abstract: Predictive applications of machine learning often rely on small (sub 1 Bn parameter) specialized models tuned to particular domains or modalities. S

MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese

Model ReleasesDGX agent

arXiv:2604.25926v1 Announce Type: new Abstract: The use of large language models (LLMs) for complex mathematical reasoning is an emergent area of research, with fast progress in methods, models, and b

Mistral AI made the TIME100 Most Influential Companies list for 2026 — and the top 10 for AI. Why we're proud: customers run frontier models…

Model ReleasesDGX agent

Mistral AI made the TIME100 Most Influential Companies list for 2026 — and the top 10 for AI. Why we're proud: customers run frontier models in production on their own terms, on their own infrastructu

MixerCA: An Efficient and Accurate Model for High-Performance Hyperspectral Image Classification

Model ReleasesDGX agent

arXiv:2604.26138v1 Announce Type: new Abstract: Over the past decade, hyperspectral image (HSI) classification has drawn considerable interest due to HSIs' ability to effectively distinguish terrestri

MoRFI: Monotonic Sparse Autoencoder Feature Identification

Model ReleasesDGX agent

arXiv:2604.26866v1 Announce Type: new Abstract: Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction. Subsequent stages of

Multimodal LLMs are not all you need for Pediatric Speech Language Pathology

Model ReleasesDGX agent

arXiv:2604.26568v1 Announce Type: new Abstract: Speech Sound Disorders (SSD) affect roughly five percent of children, yet speech-language pathologists face severe staffing shortages and unmanageable c

Naamah: A Large Scale Synthetic Sanskrit NER Corpus via DBpedia Seeding and LLM Generation

Model ReleasesDGX agent

arXiv:2604.26456v1 Announce Type: cross Abstract: The digitisation of classical Sanskrit literature is impeded by a scarcity of annotated resources, particularly for Named Entity Recognition. While re

NeuroPlastic: A Plasticity-Modulated Optimizer for Biologically Inspired Learning Dynamics

Model ReleasesDGX agent

arXiv:2604.26297v1 Announce Type: new Abstract: Optimization algorithms are fundamental to modern deep learning, yet most widely used methods rely on update rules based primarily on local gradient sta

Now available for ChatGPT accounts: Advanced Account Security, a new opt-in setting for people at higher risk of digital attacks, with stron…

Model ReleasesDGX agent

Now available for ChatGPT accounts: Advanced Account Security, a new opt-in setting for people at higher risk of digital attacks, with stronger protections including phishing-resistant sign-in and mor

NTIRE 2026 3D Restoration and Reconstruction in Real-world Adverse Conditions: RealX3D Challenge Results

Model ReleasesDGX agent

arXiv:2604.04135v2 Announce Type: replace Abstract: This paper presents a comprehensive review of the NTIRE 2026 3D Restoration and Reconstruction (3DRR) Challenge, detailing the proposed methods and

OMEGA: Optimizing Machine Learning by Evaluating Generated Algorithms

Model ReleasesDGX agent

arXiv:2604.26211v1 Announce Type: new Abstract: In order to automate AI research we introduce a full, end-to-end framework, OMEGA: Optimizing Machine learning by Evaluating Generated Algorithms, that

Omni2Sound: Towards Unified Video-Text-to-Audio Generation

Model ReleasesDGX agent

arXiv:2601.02731v3 Announce Type: replace-cross Abstract: Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers signifi

OpenAI says its models, starting with GPT-5.1, 'increasingly mentioned goblins, gremlins, and other creatures', leading to prompt instructions to mitigate it (OpenAI)

Model ReleasesDGX agent

OpenAI: OpenAI says its models, starting with GPT-5.1, “increasingly mentioned goblins, gremlins, and other creatures”, leading to prompt instructions to mitigate it — Starting with GPT-5.1, our model

OpenAI’s new security model is for ‘critical cyber defenders’ only

Model ReleasesDGX agent

OpenAI is preparing to launch a new frontier cybersecurity model, GPT-5.5-Cyber. CEO Sam Altman said the model will not be available to the general public, but will be first rolled out to a select gro

OpenMetadata maker Collate launches AI Analytics for chat-driven dashboards

Model ReleasesDGX agent

Semantic intelligence company Collate Inc. today announced the launch of Collate AI Analytics, a new chat-based tool that lets data analysts find data sources, write queries and build dashboards from

Option-Order Randomisation Reveals a Distributional Position Attractor in Prompted Sandbagging

Model ReleasesDGX agent

arXiv:2604.26206v1 Announce Type: cross Abstract: A predecessor pilot (Cacioli, 2026) found that Llama-3-8B implements prompted sandbagging as positional collapse rather than answer avoidance. However

Our evaluation of OpenAI's GPT-5.5 cyber capabilities

Model ReleasesDGX agent

Our evaluation of OpenAI's GPT-5.5 cyber capabilities The UK's AI Security Institute previously evaluated Claude Mythos: now they've evaluated GPT-5.5 for finding security vulnerability and found it t

Parameterized Quantum Circuits as Feature Maps: Representation Quality and Readout Effects in Multispectral Land-Cover Classification

Model ReleasesDGX agent

arXiv:2604.26675v1 Announce Type: cross Abstract: We investigate variational quantum classifiers (VQCs) for land-cover classification from multispectral satellite imagery, adopting a feature-map persp

PATCH: Learnable Tile-level Hybrid Sparsity for LLMs

Model ReleasesDGX agent

arXiv:2509.23410v4 Announce Type: replace-cross Abstract: Large language models (LLMs) deliver impressive performance but incur prohibitive memory and compute costs at deployment. Model pruning is an

PEOPLE ARE NOW RUNNING CLAUDE CODE WITH LOCAL AI MODELS TO AVOID API COSTS. By connecting tools like Ollama and Gemma 4, developers can buil…

Model ReleasesDGX agent

PEOPLE ARE NOW RUNNING CLAUDE CODE WITH LOCAL AI MODELS TO AVOID API COSTS. By connecting tools like Ollama and Gemma 4, developers can build apps locally with unlimited usage and no monthly billing.

Perception Test 2025: Challenge Summary and a Unified VQA Extension

Model ReleasesDGX agent

arXiv:2601.06287v2 Announce Type: replace Abstract: The Third Perception Test challenge was organised as a full-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2

Preserving Disagreement: Architectural Heterogeneity and Coherence Validation in Multi-Agent Policy Simulation

Model ReleasesDGX agent

arXiv:2604.26561v1 Announce Type: cross Abstract: Multi-agent deliberation systems using large language models (LLMs) are increasingly proposed for policy simulation, yet they suffer from artificial c

Progressive Semantic Communication for Efficient Edge-Cloud Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.26508v1 Announce Type: cross Abstract: Deploying Vision-Language Models (VLMs) on edge devices remains challenging due to their substantial computational and memory demands, which exceed th

QERNEL: a Scalable Large Electron Model

Model ReleasesDGX agent

arXiv:2604.26018v1 Announce Type: cross Abstract: We introduce QERNEL, a foundational neural wavefunction that variationally solves families of parameterized many-electron Hamiltonians and captures th

Quantum Feature Selection with Higher-Order Binary Optimization on Trapped-Ion Hardware

Model ReleasesDGX agent

arXiv:2604.26834v1 Announce Type: cross Abstract: We present a quantum feature-selection framework based on a higher-order unconstrained binary optimization (HUBO) formulation that explicitly incorpor

QYOLO: Lightweight Object Detection via Quantum Inspired Shared Channel Mixing

Model ReleasesDGX agent

arXiv:2604.26435v1 Announce Type: cross Abstract: The rapid advancement of object detection architectures has positioned single stage detectors as the dominant solution for real-time visual perception

RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments

Model ReleasesDGX agent

arXiv:2604.26067v1 Announce Type: new Abstract: We present RADIO-ViPE (Reduce All Domains Into One -- Video Pose Engine), an online semantic SLAM system that enables geometry-aware open-vocabulary gro

RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2604.26039v1 Announce Type: cross Abstract: The optimal kernel configuration for Mixture-of-Experts (MoE) inference depends on both batch size and the expert routing distribution, yet production

Random Cloud: Finding Minimal Neural Architectures Without Training

Model ReleasesDGX agent

arXiv:2604.26830v1 Announce Type: cross Abstract: I propose the Random Cloud method, a training-free approach to neural architecture search that discovers minimal feedforward network topologies throug

Reasoning Gets Harder for LLMs Inside A Dialogue

Model ReleasesDGX agent

arXiv:2603.20133v2 Announce Type: replace Abstract: Large Language Models (LLMs) achieve strong performance on many reasoning benchmarks, yet these evaluations typically focus on isolated tasks that d

ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization

Model ReleasesDGX agent

arXiv:2602.15983v2 Announce Type: replace-cross Abstract: Large language models (LLMs) can translate natural language into optimization code, but silent failures pose a critical risk: code that execut

resharing this note, find it helpful given all the great open evals work + teams building vertical agents Evals are a proxy for the behavior…

Model ReleasesDGX agent

resharing this note, find it helpful given all the great open evals work + teams building vertical agents Evals are a proxy for the behavior we want our agent to exhibit in production Model+Harness pu

Retrieval-Augmented LLMs for Evidence Localization in Clinical Trial Recruitment from Longitudinal EHR Narratives

Model ReleasesDGX agent

arXiv:2604.05190v2 Announce Type: replace-cross Abstract: Screening patients for enrollment is a well-known, labor-intensive bottleneck that leads to under-enrollment and, ultimately, trial failures.

reward-lens: A Mechanistic Interpretability Library for Reward Models

Model ReleasesDGX agent

arXiv:2604.26130v1 Announce Type: cross Abstract: Every RLHF-trained language model is shaped by a reward model, yet the mechanistic interpretability toolkit -- logit lens, direct logit attribution, a

Runpod launches Flash to bring AI inference to developers without infra overhead

Model ReleasesDGX agent

Developer-centered artificial intelligence cloud provider Runpod Inc. today announced the launch of Flash, a software development kit and platform that removes the infrastructure overhead for deployin

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment

Model ReleasesDGX agent

arXiv:2601.04389v2 Announce Type: replace-cross Abstract: Current safety evaluations of large language models (LLMs) create a dangerous illusion of universal protection by aggregating harms under gene

Say hello to the Cohere Centre! 🇨🇦 We’re proud to partner with Ottawa’s premier convention and event facility as it enters an exciting new…

Model ReleasesDGX agent

Say hello to the Cohere Centre! 🇨🇦 We’re proud to partner with Ottawa’s premier convention and event facility as it enters an exciting new chapter. The Cohere Centre will serve as a hub where leaders

SciMDR: Advancing Scientific Multimodal Document Reasoning

Model ReleasesDGX agent

arXiv:2603.12249v2 Announce Type: replace-cross Abstract: Constructing scientific multimodal document reasoning datasets for foundation model training involves an inherent trade-off among scale, faith

SEAL: Semantic-aware Single-image Sticker Personalization with a Large-scale Sticker-tag Dataset

Model ReleasesDGX agent

arXiv:2604.26883v1 Announce Type: new Abstract: Synthesizing a target concept from a single reference image is challenging in diffusion-based personalized text-to-image generation, particularly for st

Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training

Model ReleasesDGX agent

arXiv:2510.20956v2 Announce Type: replace-cross Abstract: We discover a novel and surprising phenomenon of unintentional misalignment in reasoning language models (RLMs), which we call self-jailbreaki

Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens

Model ReleasesDGX agent

arXiv:2604.26355v1 Announce Type: new Abstract: Reasoning in Large Language Models incurs significant inference-time compute, yet the token-level information structure of reasoning traces remains unde

Single unedited raw photo vs processed stack, captured from behind the moon. The color is there, just so faint it takes dozens of photos to …

Model ReleasesDGX agent

Single unedited raw photo vs processed stack, captured from behind the moon. The color is there, just so faint it takes dozens of photos to extract Lunar photography to the absolute extreme Just a hin

SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment

Model ReleasesDGX agent

arXiv:2604.25937v1 Announce Type: cross Abstract: Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professi

StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario

Model ReleasesDGX agent

arXiv:2604.26500v1 Announce Type: new Abstract: LLMs and speech assistants are increasingly used for task-oriented interactions, yet their evaluation often relies on controlled scenarios that fail to

State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading

Model ReleasesDGX agent

arXiv:2604.26614v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved impressive progress on general multimodal tasks, yet they remain brittle on dial-based measuremen

StratMem-Bench: Evaluating Strategic Memory Use in Virtual Character Conversation Beyond Factual Recall

Model ReleasesDGX agent

arXiv:2604.26243v1 Announce Type: cross Abstract: Achieving realistic human-like conversation for virtual characters requires not only a simple memorization and recall of past events, but also the str

← Previous
1…297298299300301…376
Next →