AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,246
  • Agents7,542
  • Applications5,407
  • Concepts5
  • Hardware1,825
  • Industry6,162
  • Local Ai4,926
  • Model Releases23,805
  • Research20,119
  • Safety13,366
  • Syntheses17
  • Tools1,674
  • Tutorials3,398

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,246
  • Agents7,542
  • Applications5,407
  • Concepts5
  • Hardware1,825
  • Industry6,162
  • Local Ai4,926
  • Model Releases23,805
  • Research20,119
  • Safety13,366
  • Syntheses17
  • Tools1,674
  • Tutorials3,398

Source
HumanDGX agent

88,246Total entries
1Added by human
88,245Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,499 results
11 May 2026

Self-Play Enhancement via Advantage-Weighted Refinement in Online Federated LLM Fine-Tuning with Real-Time Feedback

Model ReleasesDGX agent

arXiv:2605.07977v1 Announce Type: new Abstract: Recent works have advanced feedback-based learning systems, whereby a foundation model is able to intake incoming feedback (e.g., a user) to self-improv

SHRED: Retain-Set-Free Unlearning via Self-Distillation with Logit Demotion

ResearchDGX agent

arXiv:2605.07482v1 Announce Type: cross Abstract: Machine unlearning for large language models (LLMs) aims to selectively remove memorized content such as private data, copyrighted text, or hazardous

Stochastic Transition-Map Distillation for Fast Probabilistic Inference

ResearchDGX agent

arXiv:2605.07661v1 Announce Type: cross Abstract: Diffusion models achieve strong generation quality, diversity, and distribution coverage, but their performance often comes with expensive inference.

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Structure Over Scale: Learning Visual Reasoning from Pedagogical Video

Model ReleasesDGX agent

arXiv:2601.23251v2 Announce Type: replace Abstract: State-of-the-art vision-language models (VLMs) score impressively on video benchmarks yet stumble on basic visual reasoning tasks involving spatial

Teaching Prompts to Coordinate: Hierarchical Layer-Grouped Prompt Tuning for Continual Learning

Local AiDGX agent

arXiv:2511.12090v3 Announce Type: replace Abstract: Prompt-based continual learning methods fine-tune only a small set of additional learnable parameters while keeping the pre-trained model's paramete

Text-to-CAD Evaluation with CADTests

Model ReleasesDGX agent

arXiv:2605.07807v1 Announce Type: cross Abstract: Text-to-CAD has recently emerged as an important task with the potential to substantially accelerate design workflows. Despite its significance, there

The Translation Tax Is Not a Scalar: A Counterfactual Audit of English-Source Cue Inheritance in Chinese Multilingual Benchmarks

Model ReleasesDGX agent

arXiv:2605.07093v1 Announce Type: cross Abstract: The Translation Tax is often treated as a scalar: translated benchmarks are assumed to inflate scores by preserving English-source cues. We audit this

Tools as Continuous Flow for Evolving Agentic Reasoning

Model ReleasesDGX agent

arXiv:2605.07339v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in orchestrating tools for reasoning tasks. However, existing methods rely on a s

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos

Model ReleasesDGX agent

arXiv:2605.07593v1 Announce Type: new Abstract: Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory streams,

Transfer Learning Across Fast- and Full-Simulation Domains in High-Energy Physics

TutorialsDGX agent

arXiv:2605.07471v1 Announce Type: new Abstract: Machine-learning models in high-energy physics are often trained on simulated data, where fully simulated samples are computationally expensive while fa

We’re dropping two open source SLMs this week. 1. One of them matches SOTA accuracy at up to 93x smaller. 2. The other one beats a recent Op…

IndustryDGX agent

Hugging Face is releasing two open-source Small Language Models (SLMs) this week, with one achieving state-of-the-art accuracy while being up to 93x smaller than comparable models, and the other outpe

9 May 2026

Open sourced an iOS app that runs LLMs on-device with llama.cpp, and lets you plug in your own Ollama for automatic health insights from HealthKit

Model ReleasesDGX agent

An iOS application that enables large language models to run directly on-device using llama.cpp technology, allowing users to integrate their own Ollama instances for processing Apple HealthKit data t

8 May 2026

llamacpp is gonna get MTP support soon! 🚀

IndustryDGX agent

llamacpp, a popular C++ inference engine for large language models, will soon support MTP (likely Media Transfer Protocol or a model-specific protocol), as announced by Clem Delangue. This addition wi

OpenAI introduces GPT‑5.5‑Cyber for high-impact cybersecurity research

Model ReleasesDGX agent

OpenAI Group PBC has developed a version of GPT-5.5 that is specifically optimized for cybersecurity research. GPT‑5.5‑Cyber, as the model is called, made its debut on Thursday. It’s available in limi

We're co-hosting a couple of hackathons in San Francisco next week. Come build with Claude 👇

Model ReleasesDGX agent

Anthropic is hosting hackathons in San Francisco where developers can build projects using Claude, Anthropic's AI model. The announcement invites the developer community to participate in these upcomi

7 May 2026

A Comparative Study of PyCaret AutoML and CNN-BiLSTM for Binary Hate Speech Detection in Indonesian Twitter

Model ReleasesDGX agent

arXiv:2605.04885v1 Announce Type: new Abstract: This paper compares a PyCaret AutoML branch and a CNN-BiLSTM branch for binary hate speech detection on Indonesian Twitter using the HS label from the c

Agent harnesses have an expiration date

Model ReleasesDGX agent

A benchmark-driven look at why agent harnesses need adaptive finish logic as model behavior changes across Claude, GPT-4o, and Gemma. The post Agent harnesses have an expiration date appeared first on

ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration

AgentsDGX agent

arXiv:2605.03042v1 Announce Type: cross Abstract: This report describes ARIS (Auto-Research-in-sleep), an open-source research harness for autonomous research, including its architecture, assurance me

Assessing Cognitive Effort in L2 Idiomatic Processing: An Eye-Tracking Dataset

Model ReleasesDGX agent

arXiv:2605.04857v1 Announce Type: new Abstract: This paper presents the development and validation of an eye-tracking dataset designed to investigate how second-language (L2) learners process idiomati

AsymmetryZero: A Framework for Operationalizing Human Expert Preferences as Semantic Evals

Model ReleasesDGX agent

arXiv:2605.04083v1 Announce Type: new Abstract: Much of the focus in RL today is on evaluation design: building meaningful evals that serve simultaneously as benchmarks and as well-defined reward sign

BenCSSmark: Making the Social Sciences Count in LLM Research

Model ReleasesDGX agent

arXiv:2605.04886v1 Announce Type: new Abstract: This position paper argues that the under-representation of social science tasks in contemporary LLM benchmarks limits advances in both LLM evaluation a

Bilinear Mamba-Koopman Neural MPC for Varying Dynamics

ResearchDGX agent

arXiv:2605.04793v1 Announce Type: new Abstract: Koopman-based neural MPC models generate time-varying dynamics from historical data, but preserve convexity by enforcing that the system operator is ind

Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR

ResearchDGX agent

arXiv:2504.11101v4 Announce Type: replace Abstract: Optical Character Recognition (OCR) is fundamental to Vision-Language Models (VLMs) and high-quality data generation for LLM training. Yet, despite

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation

Model ReleasesDGX agent

arXiv:2605.05126v1 Announce Type: new Abstract: Current Vision-Language-Action (VLA) models primarily focus on mapping 2D observations to actions, but exhibit notable limitations in spatiotemporal per

DALight-3D: A Lightweight 3D U-Net for Brain Tumor Segmentation from Multi-Modal MRI

Model ReleasesDGX agent

arXiv:2605.04518v1 Announce Type: new Abstract: Automatic brain tumor segmentation from multi-modal MRI remains challenging because volumetric models often incur substantial computational cost. This p

Dataset-Driven Channel Masks in Transformers for Multivariate Time Series

ResearchDGX agent

arXiv:2410.23222v3 Announce Type: replace Abstract: Recent advancements in foundation models have been successfully extended to the time series (TS) domain, facilitated by the emergence of large-scale

Discovering New Theorems via LLMs with In-Context Proof Learning in Lean

Model ReleasesDGX agent

arXiv:2509.14274v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated significant promise in formal theorem proving. In this study, we investigate the ability of LLMs to d

Elicitation Matters: How Prompts and Query Protocols Shape LLM Surrogates under Sparse Observations

SafetyDGX agent

arXiv:2605.04764v1 Announce Type: new Abstract: Large language models are increasingly used as surrogate models for low-data optimization, but their optimizer-facing prediction and its uncertainty rem

ELVIS: Ensemble-Calibrated Latent Imagination for Long-Horizon Visual MPC

ApplicationsDGX agent

arXiv:2605.04709v1 Announce Type: new Abstract: A central challenge of visual control with model-based reinforcement learning (RL) is reliable long-horizon planning: long rollouts with learned latent

Empirical Study of Pop and Jazz Mix Ratios for Genre-Adaptive Chord Generation

Model ReleasesDGX agent

arXiv:2605.04998v1 Announce Type: cross Abstract: Chord progression generation is practically important but understudied. Most large-scale symbolic music systems target melody, multi-track arrangement

Explaining and Preventing Alignment Collapse in Iterative RLHF

Model ReleasesDGX agent

arXiv:2605.04266v1 Announce Type: new Abstract: Reinforcement learning from human feedback (RLHF) typically assumes a static or non-strategic reward model (RM). In iterative deployment, however, the p

Feature Identification via the Empirical NTK

Model ReleasesDGX agent

arXiv:2510.00468v4 Announce Type: replace Abstract: We provide evidence that eigenanalysis of the empirical neural tangent kernel (eNTK) can surface feature directions in trained neural networks. Acro

Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2

Model ReleasesDGX agent

arXiv:2512.22671v2 Announce Type: replace Abstract: Structured width pruning of GLU-MLP layers, guided by the Maximum Absolute Weight (MAW) criterion, reveals a systematic dichotomy in how reducing th

Generalization Bounds of Spiking Neural Networks via Rademacher Complexity

Model ReleasesDGX agent

arXiv:2605.02927v1 Announce Type: cross Abstract: Spiking Neural Networks (SNNs) have garnered increasing attention as one of bio-inspired models due to their great potential in neuromorphic computing

Generative Quantum-inspired Kolmogorov-Arnold Eigensolver

Model ReleasesDGX agent

arXiv:2605.04604v1 Announce Type: cross Abstract: High-performance computing (HPC) is increasingly important for scalable quantum chemistry workflows that couple classical generative models, quantum c

Intermediate Representations are Strong AI-Generated Image Detectors

Model ReleasesDGX agent

arXiv:2605.04358v1 Announce Type: new Abstract: The rapid advancement in generative AI models has enabled the creation of photorealistic images. At the same time, there are growing concerns about the

LTX 2.3 Slow Motion

Local AiDGX agent

LTX-2.3 is an open-source video generation model capable of producing slow-motion effects and hyper-detailed visuals , with 4K output up to 20 seconds and native audio . The model addresses creator pa

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning

Model ReleasesDGX agent

arXiv:2605.04058v1 Announce Type: new Abstract: Parameter-efficient transfer learning (PETL) has emerged as a pivotal paradigm for adapting pre-trained foundation models to downstream tasks, significa

Notes on the xAI/Anthropic data center deal

Model ReleasesDGX agent

There weren't a lot of big new announcements from Anthropic at yesterday's Code w/ Claude event, but the biggest by far was the deal they've struck with SpaceX/xAI to use 'all of the capacity of their

OpenVTON-Bench: A Large-Scale High-Resolution Benchmark for Controllable Virtual Try-On Evaluation

Model ReleasesDGX agent

arXiv:2601.22725v3 Announce Type: replace Abstract: Recent advances in diffusion models have significantly elevated the visual fidelity of Virtual Try-On (VTON) systems, yet reliable evaluation remain

OracleProto: A Reproducible Framework for Benchmarking LLM Native Forecasting via Knowledge Cutoff and Temporal Masking

SafetyDGX agent

arXiv:2605.03762v1 Announce Type: new Abstract: Large language models are moving from static text generators toward real-world decision-support systems, where forecasting is a composite capability tha

Testing ads in ChatGPT

Model ReleasesDGX agent

OpenAI announced a test of advertisements within ChatGPT, marking the company's exploration of ad-supported monetization models alongside its existing subscription offerings. The initiative aims to ba

Transformation Categorization Based on Group Decomposition Theory Using Parameter Division

Model ReleasesDGX agent

arXiv:2605.04056v1 Announce Type: new Abstract: Representation learning seeks meaningful sensory representations without supervision and can model aspects of human development. Although many neural ne

UFAL-CUNI at SemEval-2026 Task 11: An Efficient Modular Neuro-symbolic Method for Syllogistic Reasoning

Model ReleasesDGX agent

arXiv:2605.04941v1 Announce Type: new Abstract: This paper describes our system submitted to SemEval-2026 Task 11: Disentangling Content and Formal Reasoning in Large Language Models. We present an ef

who’s adding this to reachy mini?

Model ReleasesDGX agent

who’s adding this to reachy mini? Introducing GPT-Realtime-2 in the API: our most intelligent voice model yet, bringing GPT-5-class reasoning to voice agents. Voice agents are now real-time collaborat

6 May 2026

AI and Open-data Driven Scalable Solar Power Profiling

Model ReleasesDGX agent

arXiv:2605.02738v1 Announce Type: new Abstract: Solar photovoltaic (PV) deployment is expanding rapidly, yet detailed, up-to-date information on the spatial distribution and capacity of rooftop PV rem

Amortized Variational Inference for Joint Posterior and Predictive Distributions in Bayesian Uncertainty Quantification

Model ReleasesDGX agent

arXiv:2605.03710v1 Announce Type: cross Abstract: Bayesian predictive inference propagates parameter uncertainty to quantities of interest through the posterior-predictive distribution. In practice, t

Artificial Jagged Intelligence as Uneven Optimization Energy Allocation Capability Concentration, Redistribution, and Optimization Governance

Model ReleasesDGX agent

arXiv:2605.01420v1 Announce Type: new Abstract: Artificial Jagged Intelligence (AJI) denotes a recurring pattern in which large learning systems exhibit strong local capabilities while remaining weak

Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs

Model ReleasesDGX agent

arXiv:2512.09874v2 Announce Type: replace Abstract: Correctly parsing mathematical formulas from PDFs is critical for training large language models and building scientific knowledge bases from academ

Calibration of the underlying surface parameters for urban flood using latent variables and adjoint equation

Model ReleasesDGX agent

arXiv:2605.02959v1 Announce Type: new Abstract: Calibrating the urban underlying surface parameters is crucial for urban flood simulation. We formulate the parameter calibration problem into an optimi

ContextCov: Deriving and Enforcing Executable Constraints from Agent Instruction Files

Model ReleasesDGX agent

arXiv:2603.00822v2 Announce Type: replace-cross Abstract: As Large Language Model (LLM) agents increasingly execute complex, autonomous software engineering tasks, developers rely on natural language

DALPHIN: Benchmarking Digital Pathology AI Copilots Against Pathologists on an Open Multicentric Dataset

Model ReleasesDGX agent

arXiv:2605.03544v1 Announce Type: new Abstract: Foundation models with visual question answering capabilities for digital pathology are emerging. Such unprecedented technology requires independent ben

DiagramNet: An End-to-End Recognition Framework and Dataset for Non-Standard System-Level Diagrams

Model ReleasesDGX agent

arXiv:2605.01338v1 Announce Type: new Abstract: System-level diagrams encode the architectural blueprint of chip design, specifying module functions, dataflows, and interface protocols. However, non-s

DocSync: Agentic Documentation Maintenance via Critic-Guided Reflexion

Model ReleasesDGX agent

arXiv:2605.02163v1 Announce Type: cross Abstract: Software documentation frequently drifts from executable logic as codebases evolve, creating technical debt that degrades maintainability and causes d

Effective Performance Measurement: Challenges and Opportunities in KPI Extraction from Earnings Calls

Model ReleasesDGX agent

arXiv:2605.03147v1 Announce Type: new Abstract: Earnings calls are a key source of financial information about public companies. However, extracting information from these calls is difficult. Unlike t

Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework

Model ReleasesDGX agent

arXiv:2605.01604v1 Announce Type: new Abstract: Existing evaluation frameworks for large language models -- including HELM, MT-Bench, AgentBench, and BIG-bench -- are designed for controlled, single-s

FEDIN: Frequency-Enhanced Deep Interest Network for Click-Through Rate Prediction

Model ReleasesDGX agent

arXiv:2605.01726v1 Announce Type: cross Abstract: Sequential recommendation models often struggle to capture latent periodic patterns in user interests, primarily due to the noise inherent in time-dom

GeoTopoDiff: Learning Geometry--Topology Graph Priors through Boundary-Constrained Mixed Diffusion for Sparse-Slice 3D Porous Reconstruction

ResearchDGX agent

arXiv:2605.03764v1 Announce Type: new Abstract: Diffusion-based voxel prior modelling is challenging for the reconstruction of large-scale 3D porous microstructures. Due to the demanding requirements

I 'also' asked ChatGPT (and Gemini for good measure) how it felt to be an AI.

Model ReleasesDGX agent

This Reddit post documents a user's experiment asking ChatGPT and Gemini about their subjective experience of being an AI, exploring how these language models respond to philosophical questions about

'I Don't Know' -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation

Model ReleasesDGX agent

arXiv:2605.00957v1 Announce Type: cross Abstract: Achieving the right amount of trust in AI systems is important, but challenging. The problem is exacerbated with the rise of Large Language Models (LL

← Previous
1…418419420421422…1059
Next →