AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,553 results
4 May 2026

Caracal: Causal Architecture via Spectral Mixing

Model ReleasesDGX agent

arXiv:2605.00292v1 Announce Type: new Abstract: The scalability of Large Language Models to long sequences is hindered by the quadratic cost of attention and the limitations of positional encodings. T

claudely: launch Claude Code against Local LLM provider like LM Studio / Ollama / llama.cpp without trashing your real claude config

Model ReleasesDGX agent

claudely is a tool that enables users to run Claude Code against local LLM providers such as LM Studio, Ollama, or llama.cpp while preserving their existing Claude configuration. The tool allows devel

CMTA: Leveraging Cross-Modal Temporal Artifacts for Generalizable AI-Generated Video Detection

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.00630v1 Announce Type: new Abstract: The proliferation of advanced AI video synthesis techniques poses an unprecedented challenge to digital video authenticity. Existing AI-generated video

CompleteRXN: Toward Completing Open Chemical Reaction Databases

Model ReleasesDGX agent

arXiv:2605.00222v1 Announce Type: new Abstract: Chemical reaction datasets such as USPTO suffer from substantial incompleteness, frequently missing byproducts, co-reactants, and stoichiometric coeffic

Concolic Testing on Individual Fairness of Neural Network Models

Model ReleasesDGX agent

arXiv:2509.06864v2 Announce Type: replace Abstract: This paper introduces PyFair, a formal framework for evaluating and verifying individual fairness of Deep Neural Networks (DNNs). By adapting the co

ControBench: An Interaction-Aware Benchmark for Controversial Discourse Analysis on Social Networks

Model ReleasesDGX agent

arXiv:2605.00513v1 Announce Type: new Abstract: Understanding how people argue across ideological divides online is important for studying political polarization, misinformation, and content moderatio

CURE-OOD: Benchmarking Out-of-Distribution Detection for Survival Prediction

Model ReleasesDGX agent

arXiv:2605.00350v1 Announce Type: new Abstract: ``How long can I live and remain free of cancer?'' is often the first question a patient asks after receiving a cancer diagnosis and treatment. Accurate

deepagents-cli is quietly becoming the best place to start coding with open weight models. we've been investing heavily in making it a harne…

Model ReleasesDGX agent

deepagents-cli is quietly becoming the best place to start coding with open weight models. we've been investing heavily in making it a harness that's truly model-agnostic, without compromising perform

Deepfakes: we need to re-think the concept of 'real' images

Model ReleasesDGX agent

arXiv:2509.21864v2 Announce Type: replace Abstract: The wide availability and low usability barrier of modern image generation models has triggered the reasonable fear of criminal misconduct and negat

Deepseek V4 works more thoroughly than other open source models: It writes its own tests and performs extensive validation. This leads to be…

Model ReleasesDGX agent

Deepseek V4 works more thoroughly than other open source models: It writes its own tests and performs extensive validation. This leads to better performance, but also cases of the model being overconf

Differentiable Autoencoding Neural Operator for Interpretable and Integrable Latent Space Modeling

Model ReleasesDGX agent

arXiv:2510.00233v2 Announce Type: replace Abstract: Scientific machine learning has enabled the extraction of physical insights and data-driven modeling of high-dimensional spatiotemporal data, yet ac

Do Open-Loop Metrics Predict Closed-Loop Driving? A Cross-Benchmark Correlation Study of NAVSIM and Bench2Drive

Model ReleasesDGX agent

arXiv:2605.00066v1 Announce Type: new Abstract: Open-loop evaluation offers fast, reproducible assessment of autonomous driving planners, but its ability to predict real closed-loop driving performanc

Documentation https://docs.ollama.com/integrations/claude-desktop

Model ReleasesDGX agent

This documentation page describes how to integrate Ollama with Claude Desktop, enabling users to run local language models through the Anthropic Claude interface. The integration allows Claude Desktop

Driving with A Thousand Faces: A Benchmark for Closed-Loop Personalized End-to-End Autonomous Driving

Model ReleasesDGX agent

arXiv:2602.18757v2 Announce Type: replace Abstract: Human driving behavior is inherently diverse, yet most end-to-end autonomous driving (E2E-AD) systems learn a single average driving style, neglecti

Dynamics-Encoded Deep Learning for Robust System Identification and Parameter Estimation

Model ReleasesDGX agent

arXiv:2410.04299v2 Announce Type: replace Abstract: Incorporating a priori physics knowledge into machine learning leads to more robust and interpretable algorithms. In this work, we combine deep lear

Efficient Spatio-Temporal Vegetation Pixel Classification with Vision Transformers

Model ReleasesDGX agent

arXiv:2605.00296v1 Announce Type: new Abstract: Plant phenology-the study of recurrent life cycle events-is essential for understanding ecosystem dynamics and their responses to climate change impacts

Evaluating the Architectural Reasoning Capabilities of LLM Provers via the Obfuscated Natural Number Game

Model ReleasesDGX agent

arXiv:2605.00677v1 Announce Type: new Abstract: While Large Language Models have achieved notable success on formal mathematics benchmarks such as MiniF2F, it remains unclear whether these results ste

Event-based Civil Infrastructure Visual Defect Detection: ev-CIVIL Dataset and Benchmark

Model ReleasesDGX agent

arXiv:2504.05679v2 Announce Type: replace Abstract: Small unmanned aerial vehicle (UAV)-based visual inspections are a more efficient alternative to manual methods for examining civil structural defec

Ex-iRobot CEO Colin Angle launches Familiar Machines & Magic and unveils Familiar, a dog-like, 'emotionally intelligent' robot that reacts to owner's feelings (Christopher Mims/Wall Street Journal)

Model ReleasesDGX agent

Christopher Mims / Wall Street Journal: Ex-iRobot CEO Colin Angle launches Familiar Machines & Magic and unveils Familiar, a dog-like, “emotionally intelligent” robot that reacts to owner's feelings —

ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation

Model ReleasesDGX agent

arXiv:2507.14201v3 Announce Type: replace-cross Abstract: We present ExCyTIn-Bench, the first benchmark to Evaluate an LLM agent X on the task of Cyber Threat Investigation through security questions

Exploring the System 1 Thinking Capability of Large Reasoning Models

Model ReleasesDGX agent

arXiv:2504.10368v4 Announce Type: replace Abstract: This paper explores the system 1 thinking capability of Large Reasoning Models (LRMs), the intuitive ability to respond efficiently with minimal tok

Faithful Extreme Image Rescaling with Learnable Reversible Transformation and Semantic Priors

Model ReleasesDGX agent

arXiv:2605.00605v1 Announce Type: new Abstract: Most recent extreme rescaling methods struggle to preserve semantically consistent structures and produce realistic details, due to the severely ill-pos

FedACT: Concurrent Federated Intelligence across Heterogeneous Data Sources

Model ReleasesDGX agent

arXiv:2605.00011v1 Announce Type: new Abstract: Federated Learning (FL) enables collaborative intelligence across decentralized data source devices in a privacy-preserving way. While substantial resea

FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios

Model ReleasesDGX agent

arXiv:2605.00706v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied in financial scenarios. However, they may produce harmful outputs, including facilitating illegal

Firestore at Next '26: Unlock agentic development, search and MongoDB compatibility

Model ReleasesDGX agent

In the era of AI agents, the distance between a big idea and a working application has never been shorter. As we lean more heavily on agents to help us build applications, a critical question remains:

FollowTable: A Benchmark for Instruction-Following Table Retrieval

Model ReleasesDGX agent

arXiv:2605.00400v1 Announce Type: cross Abstract: Table Retrieval (TR) has traditionally been formulated as an ad-hoc retrieval problem, where relevance is primarily determined by topical semantic sim

Foresight Arena: An On-Chain Benchmark for Evaluating AI Forecasting Agents

Model ReleasesDGX agent

arXiv:2605.00420v1 Announce Type: cross Abstract: Evaluating the true forecasting ability of AI agents requires environments resistant to overfitting, free from centralized trust, and grounded in ince

From Backward Spreading to Forward Replay: Revisiting Target Construction in LLM Parameter Editing

Model ReleasesDGX agent

arXiv:2605.00358v1 Announce Type: new Abstract: LLM parameter editing methods commonly rely on computing an ideal target hidden-state at a target layer (referred as anchor point) and distributing the

From Images2Mesh: A 3D Surface Reconstruction Pipeline for Non-Cooperative Space Objects

Model ReleasesDGX agent

arXiv:2605.00147v1 Announce Type: new Abstract: On-orbit inspection imagery is crucial as it enables characterization of non-cooperative resident space objects, providing the geometry and structural c

From Prediction to Practice: A Task-Aware Evaluation Framework for Blood Glucose Forecasting

Model ReleasesDGX agent

arXiv:2605.00645v1 Announce Type: new Abstract: Clinical time-series forecasting is increasingly studied for decision support, yet standard aggregate metrics can obscure whether a model is actually us

GitHub: ComfyUI SenseNova U1 Released – Anyone Got It Working Yet for ComfyUI?

Model ReleasesDGX agent

SenseNova U1 is a new series of native multimodal models that unifies multimodal understanding, reasoning, and generation within a monolithic architecture, marking a fundamental paradigm shift from mo

Granite 4.1 3B SVG Pelican Gallery

Model ReleasesDGX agent

Granite 4.1 3B SVG Pelican Gallery IBM released their Granite 4.1 family of LLMs a few days ago. They're Apache 2.0 licensed and come in 3B, 8B and 30B sizes. Granite 4.1 LLMs: How They’re Built by Gr

Grok 4.3 just built this entire game with just a single prompt It has the fastest output token speed and outranks Claude Sonnet 4.6 Max on A…

Model ReleasesDGX agent

Grok 4.3 just built this entire game with just a single prompt It has the fastest output token speed and outranks Claude Sonnet 4.6 Max on Artificial Analysis I built this using the xAI API in Kilo Co

High-Probability Convergence in Decentralized Stochastic Optimization with Gradient Tracking

Model ReleasesDGX agent

arXiv:2605.00281v1 Announce Type: new Abstract: We study high-probability (HP) convergence guarantees in decentralized stochastic optimization, where multiple agents collaborate to jointly train a mod

How Frontier LLMs Adapt to Neurodivergence Context: A Measurement Framework for Surface vs. Structural Change in System-Prompted Responses

Model ReleasesDGX agent

arXiv:2605.00113v1 Announce Type: new Abstract: We examine if frontier chat-based large language models (LLMs) adjust their outputs based on neurodivergence (ND) context in system prompts and describe

How I built a free, local AI powerhouse in 10 days (Ollama + Gemma 4 + Claude Cowork 3P + Browserless)

Model ReleasesDGX agent

This post documents a 10-day project to build a local AI system using open-source tools and models, specifically combining Ollama (a local LLM framework), Gemma 4 (a language model), Claude Cowork 3P,

How OpenAI delivers low-latency voice AI at scale

Model ReleasesDGX agent

OpenAI describes its technical approach to delivering real-time voice AI services with minimal latency across large user bases, likely covering infrastructure optimization, model serving strategies, a

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks

Model ReleasesDGX agent

arXiv:2507.01955v3 Announce Type: replace Abstract: Multimodal foundation models (MFMs), such as GPT-4o, have recently made remarkable progress. However, their detailed visual understanding beyond que

Hypergraph and Latent ODE Learning for Multimodal Root Cause Localization in Microservices

Model ReleasesDGX agent

arXiv:2605.00351v1 Announce Type: new Abstract: Root cause localization in cloud native microservice systems requires modeling complex service dependencies, irregular temporal dynamics, and heterogene

I had Claude Code for web build me this WebAssembly playground for trying out the new Redis array commands https://tools.simonwillison.net/r…

Model ReleasesDGX agent

I had Claude Code for web build me this WebAssembly playground for trying out the new Redis array commands https://tools.simonwillison.net/redis-array More notes here: https://simonwillison.net/2026/M

I think the fact that GPT-4o and Llama 3.3-80B did no significant harm is just as important as whether AI helped. If older (less accurate & …

Model ReleasesDGX agent

I think the fact that GPT-4o and Llama 3.3-80B did no significant harm is just as important as whether AI helped. If older (less accurate & more sycophantic) chatbots essentially did nothing for peopl

InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information

Model ReleasesDGX agent

arXiv:2508.07630v2 Announce Type: replace Abstract: We introduce InterChart, a diagnostic benchmark that evaluates how well vision-language models (VLMs) reason across multiple related charts, a task

Introducing nanowhale 🐳! A tiny DeepSeek model fully pretrained by an agent. Inspired by @karpathy's nanochat, we gave ml-intern the task o…

Model ReleasesDGX agent

Introducing nanowhale 🐳! A tiny DeepSeek model fully pretrained by an agent. Inspired by @karpathy's nanochat, we gave ml-intern the task of training a tiny MoE with all the architectural advancements

Introducing WARM-VR: Benchmark Dataset for Multimodal Wearable Affect Recognition in Virtual Reality

Model ReleasesDGX agent

arXiv:2605.00184v1 Announce Type: new Abstract: With the growing integration of human-computer interaction into everyday life, advances in machine learning have enabled systems to better perceive and

Jailbreaking Vision-Language Models Through the Visual Modality

Model ReleasesDGX agent

arXiv:2605.00583v1 Announce Type: new Abstract: The visual modality of vision-language models (VLMs) is an underexplored attack surface for bypassing safety alignment. We introduce four jailbreak atta

Jailbroken Frontier Models Retain Their Capabilities

Model ReleasesDGX agent

arXiv:2605.00267v1 Announce Type: new Abstract: As language model safeguards become more robust, attackers are pushed toward developing increasingly complex jailbreaks. Prior work has found that this

Knowing When to Defer: Selective Prediction for Responsible Knowledge Tracing

Model ReleasesDGX agent

arXiv:2509.21514v3 Announce Type: replace-cross Abstract: Research on Knowledge Tracing (KT) models traditionally focuses on improving predictive accuracy. However, responsible real-world deployment r

Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis

Model ReleasesDGX agent

arXiv:2605.00448v1 Announce Type: new Abstract: The deployment of artificial intelligence in medical imaging is hindered by high computational complexity and resource-intensive processing of volumetri

Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation

Model ReleasesDGX agent

arXiv:2510.19897v2 Announce Type: replace Abstract: We investigate how agents built on pretrained large language models (LLMs) can learn target classification functions from labeled examples without p

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation

Model ReleasesDGX agent

arXiv:2605.00051v1 Announce Type: new Abstract: Anticipating traffic accidents is a critical yet unresolved problem for autonomous driving, hindered by the inherent complexity of modeling interactions

Learning Locally, Revising Globally: Global Reviser for Federated Learning with Noisy Labels

Model ReleasesDGX agent

arXiv:2412.00452v2 Announce Type: replace-cross Abstract: Conventioanl federated learning (FL) heavily depends on high-quality labels, which are often impractical in the real world, leading to the fed

Lightweight Domain Adaptation of a Large Language Model for Legal Assistance in the Indian Context

Model ReleasesDGX agent

arXiv:2505.22003v2 Announce Type: replace Abstract: In India, access to legal assistance for the general public has been observed to have a critical gap, as many citizens are not able to take full adv

LLM-Oriented Information Retrieval: A Denoising-First Perspective

Model ReleasesDGX agent

arXiv:2605.00505v1 Announce Type: cross Abstract: Modern information retrieval (IR) is no longer consumed primarily by humans but increasingly by large language models (LLMs) via retrieval-augmented g

LWiAI Podcast #243 - GPT 5.5, DeepSeek V4, AI safety sabotage

Model ReleasesDGX agent

This podcast episode from Last Week in AI discusses recent developments in large language models, including updates on GPT 5.5 and DeepSeek V4, while also covering concerns about potential sabotage or

M-CaStLe: Uncovering Local Causal Structures in Multivariate Space-Time Gridded Data

Model ReleasesDGX agent

arXiv:2605.00398v1 Announce Type: new Abstract: Causal graph discovery for space-time systems is challenging in high-dimensional gridded data, which often has many more grid cells than temporal observ

Make Your LVLM KV Cache More Lightweight

Model ReleasesDGX agent

arXiv:2605.00789v1 Announce Type: new Abstract: Key-Value (KV) cache has become a de facto component of modern Large Vision-Language Models (LVLMs) for inference. While it enhances decoding efficiency

May 5 is the GPT-5.5 launch celebration in San Francisco and the Claude Finance Briefing in New York. Real opposite valence events on opposi…

Model ReleasesDGX agent

I cannot provide a summary for this entry as the post content appears incomplete or corrupted in the provided information. The title is cut off mid-sentence and doesn't clearly convey the full topic.

MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems

Model ReleasesDGX agent

arXiv:2510.17281v5 Announce Type: replace Abstract: Scaling up data, parameters, and test-time computation has been the mainstream methods to improve LLM systems (LLMsys), but their upper bounds are a

Minimizing Human Intervention in Online Classification

Model ReleasesDGX agent

arXiv:2510.23557v2 Announce Type: replace-cross Abstract: Training or fine-tuning large language model (LLM)-based systems often requires costly human feedback, yet there is limited understanding of h

ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering

Model ReleasesDGX agent

arXiv:2505.23723v2 Announce Type: replace Abstract: The emergence of large language model (LLM)-based agents has significantly advanced the development of autonomous machine learning (ML) engineering.

← Previous
1…291292293294295…376
Next →