AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

86,542Total entries
1Added by human
86,541Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,103 results
23 Jul 2026

HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

Model ReleasesDGX agent

arXiv:2607.20219v1 Announce Type: new Abstract: Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing

Kwaipilot/KAT-Coder-V2.5-Dev · Hugging Face

Model ReleasesDGX agent

from kwaipilot: Following the release of KAT-Coder-V2.5 in July, we are pleased to release the open-weight version KAT-Coder-V2.5-Dev, an MOE model with a total parameter count of 35B and 3B activated

Multi-stage Dynamic Selection for Cross-Project Defect Prediction

Model ReleasesDGX agent

arXiv:2607.20151v1 Announce Type: cross Abstract: Cross-Project Defect Prediction (CPDP) involves building models using data from external projects, called training projects, to predict modules from t

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Pixel-Space Diffusion Transformers

ResearchDGX agent

arXiv:2607.17585v2 Announce Type: replace Abstract: Latent diffusion models (LDMs) enable efficient high-resolution image synthesis by denoising in a VAE-compressed latent space. However, fixed visual

Running Qwen 3.6 35B MoE (Q4_K_M) on a Zeus (Xiaomi 12 Pro, 12GB RAM)

Model ReleasesDGX agent

Shoutout to this awesome guy - https://www.reddit.com/r/LLM/s/IDUyU3v9ap Thanks to his project, BigMoeOnEdge https://github.com/Helldez/BigMoeOnEdge, I managed to successfully run a 35B MoE model on j

STN-TGAT: Top-K Portfolio Construction via Prior-Guided Graph Attention with Learnable Soft-Threshold Sparsification

Model ReleasesDGX agent

arXiv:2607.19385v1 Announce Type: new Abstract: This paper tackles the problem of stock ranking and portfolio construction under realistic investment settings by jointly modeling temporal dynamics and

The Blessing of Dimensionality: How Near-Orthogonality in High-Dimensional Spaces Explains Temporal Portability

Model ReleasesDGX agent

arXiv:2607.20301v1 Announce Type: cross Abstract: Fine-tuning has been widely used to adapt large language models (LLMs) for domain-specific tasks. Parameter efficient fine-tuning (PEFT) methods such

The production platform for open-weight AI inference

ApplicationsDGX agent

OpenAI has updated its inference platform to give users full control over performance, cost, and quality without building their own stack—models go live in minutes and support multiple deployments beh

When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language Summarization

Model ReleasesDGX agent

arXiv:2607.19956v1 Announce Type: cross Abstract: Knowledge distillation (KD) is a standard approach for compressing sequence-to-sequence models, but its per-sample effects are rarely examined. On the

22 Jul 2026

We benchmarked Gemini 3.6 Flash and Gemini 3.5 Flash Lite on document understanding. We compared against their prior versions - Gemini 3.5 F…

Model ReleasesDGX agent

We benchmarked Gemini 3.6 Flash and Gemini 3.5 Flash Lite on document understanding. We compared against their prior versions - Gemini 3.5 Flash and Gemini 3.1 Flash Lite. 1️⃣ Gemini 3.6 Flash has rou

21 Jul 2026

Kimi K3 is second only to Fable 5 on AA-Briefcase, our agentic knowledge work benchmark, but costs more than Opus 4.8 to run while averaging…

Model ReleasesDGX agent

Kimi K3 is second only to Fable 5 on AA-Briefcase, our agentic knowledge work benchmark, but costs more than Opus 4.8 to run while averaging nearly an hour per task Last week @Kimi_Moonshot released K

this is a good summary

Model ReleasesDGX agent

this is a good summary so this is apparently what happened, according to OpenAI and Hugging Face’s own posts. wild. tl;dr: • OpenAI cyber eval – GPT-5.6 Sol and a more capable pre-release model ran Ex

16 Jul 2026

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

Model ReleasesDGX agent

arXiv:2607.13705v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluat

HEDGEHOG: Hierarchical Evaluation of Drug Generators Through Rigorous Filtration

Model ReleasesDGX agent

arXiv:2607.13155v1 Announce Type: new Abstract: Generative molecular models can support early drug discovery by proposing new candidate compounds de novo. In practice, useful candidates must balance t

Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution

Model ReleasesDGX agent

arXiv:2607.13069v1 Announce Type: new Abstract: Large language models produce chain-of-thought (CoT) reasoning that appears logically sound yet may not genuinely depend on its stated premises. We intr

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

Model ReleasesDGX agent

arXiv:2607.09142v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clini

SPINE: Bridging the Cyber-Physical Gap with Agentic AI

Model ReleasesDGX agent

arXiv:2607.13049v1 Announce Type: new Abstract: Foundation models have given robots a sophisticated brain for complex decision-making, yet deploying that intelligence into a physical platform still de

Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems

SafetyDGX agent

arXiv:2607.13048v1 Announce Type: cross Abstract: Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) that provide rich semantic understanding at

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs

Model ReleasesDGX agent

arXiv:2602.17659v2 Announce Type: replace Abstract: Vision-Language-Action models (VLAs) promise to ground language instructions in robot control, yet in practice often fail to faithfully follow langu

15 Jul 2026

1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World

Model ReleasesDGX agent

arXiv:2602.18548v2 Announce Type: replace-cross Abstract: Design-to-code translates high-fidelity UI designs into executable front-end implementations, but progress remains hard to compare due to inco

A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs

Model ReleasesDGX agent

arXiv:2607.12550v1 Announce Type: cross Abstract: The key-value (KV) cache has become the dominant memory cost of transformer inference. It grows with batch size, context length, and depth, and at lon

CANDI: Contextual Alignment for Niche Domains Question Answering

Model ReleasesDGX agent

arXiv:2607.11891v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) in specialized domains like medical diagnostics and financial advisory necessitates evaluating capabili

Controllable Generation of Diverse Dermatological Imagery for Fair and Efficient Malignancy Classification

Model ReleasesDGX agent

arXiv:2607.12987v1 Announce Type: new Abstract: Accurate dermatological diagnosis naturally necessitates equitable performance across diverse populations, yet a systematic lack of expertly annotated i

How to Analyze and Govern Gemini Enterprise App Usage at Scale with BigQuery

Model ReleasesDGX agent

Deploying the Gemini Enterprise app across an organization marks a transformative leap forward in workforce productivity, providing employees with an amazing, high-performance suite of agentic AI tool

Label-Decoupled Style Augmentation for Domain Generalization in Multi-Label Remote Sensing Scene Classification

Model ReleasesDGX agent

arXiv:2607.12704v1 Announce Type: new Abstract: Multi-label classification assigns several co-occurring labels to each aerial scene, yet deployed models often encounter data distributions different fr

Language Identification with Succinct Machine-Independent Traces

TutorialsDGX agent

arXiv:2607.12443v1 Announce Type: new Abstract: Motivated by the power of large language models, there has been renewed interest in the Gold-Angluin model of language identification in the limit, with

LLM Judges Can Be Too Generous When There Is No Reference Answer

ResearchDGX agent

arXiv:2607.12885v1 Announce Type: new Abstract: LLM judges are increasingly being used to evaluate open-ended model responses, often in no-reference settings where a ground-truth answer is unavailable

More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization

Local AiDGX agent

arXiv:2607.12429v1 Announce Type: new Abstract: Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same sc

Predict the Retrieval! Test time adaptation for Retrieval Augmented Generation

Model ReleasesDGX agent

arXiv:2601.11443v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) has emerged as a powerful approach for enhancing large language models' question-answering capabilities through

RecRec: Recursive Refinement for Sequential Recommendation

Model ReleasesDGX agent

arXiv:2607.10541v2 Announce Type: replace-cross Abstract: Sequential recommender systems typically infer user preferences through single-pass encoding of interaction histories without iterative refine

Self-Evolving In-Context Learning for Direct Pilot-to-Beamformer Design in MU-MISO Systems

Model ReleasesDGX agent

arXiv:2607.11970v1 Announce Type: cross Abstract: We develop an enhanced in-context learning (ICL) framework to improve the performance of pilot-based beamforming in multi-user multiple-input single-o

The Spectrum Is Not Enough: When Context Helps Time-Series Forecasting

ResearchDGX agent

arXiv:2607.13006v1 Announce Type: new Abstract: A growing family of indices scores how predictable a series is from its spectrum. Practitioners increasingly read these scores as answering a different

TRACE: An Operational Reasoning Schema for Auditable Agentic Commitments

Model ReleasesDGX agent

arXiv:2607.12480v1 Announce Type: new Abstract: This paper defines TRACE (Typed Reasoning And Commitment Evidence): a typed, versioned schema for recording reasoning traces, a reference procedure for

14 Jul 2026

75.4% SWE Bench Verified / 53.9% SWE Bench Pro on 1 bit quantisation is 🤪 This is in line with my expectations & you can expect even lower …

Model ReleasesDGX agent

75.4% SWE Bench Verified / 53.9% SWE Bench Pro on 1 bit quantisation is 🤪 This is in line with my expectations & you can expect even lower drop off with NVP4 base trained models - why not run everythi

10 Jul 2026

Collate: Collaborative Neural Network Learning for Latency-Critical Edge Systems

Local AiDGX agent

arXiv:2607.08013v1 Announce Type: new Abstract: Federated Learning (FL) empowers multiple clients to collaboratively learn a model, enlarging the training data of each client for high accuracy while p

Dive Into the Implicit Biases of Low-rank Vision-language Alignment

Model ReleasesDGX agent

arXiv:2607.08194v1 Announce Type: new Abstract: Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requi

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs

ApplicationsDGX agent

arXiv:2607.08009v1 Announce Type: new Abstract: We introduce a Bloom-aligned framework for measuring educational control in Large Language Models (LLMs): the ability to preserve a task's instructional

Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos

Model ReleasesDGX agent

arXiv:2512.01803v3 Announce Type: replace Abstract: Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elu

Physics-Informed Machine Learning Under Small-Data Constraints: Lessons from Abrasive Waterjet Milling

Model ReleasesDGX agent

arXiv:2607.07863v1 Announce Type: new Abstract: In physically dominated machining processes, experimental datasets are small, expensive, and material-specific; in this regime, data curation, evaluatio

RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments

Model ReleasesDGX agent

arXiv:2603.16453v3 Announce Type: replace Abstract: Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in d

The Phasor Transformer: Resolving Attention Bottlenecks on the Unit Circle

Model ReleasesDGX agent

arXiv:2603.17433v2 Announce Type: replace-cross Abstract: Transformer models have redefined sequence learning, yet dot-product self-attention introduces a quadratic token-mixing bottleneck for long-co

When Does Continual Learning Require Learning

AgentsDGX agent

arXiv:2607.07847v1 Announce Type: new Abstract: As large language models (LLMs) become increasingly capable, the next question is how can we enable models to continually learn? Today, the field largel

Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA

Model ReleasesDGX agent

arXiv:2607.08054v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly trusted to draft the artifacts of safety analysis such as, losses, hazards, Unsafe Control Actions (UCAs

9 Jul 2026

Autopilot Clusters with GKE managed DRANET: GPUs and TPUs

Model ReleasesDGX agent

Google Kubernetes Engine (GKE) managed DRANET supports both GPUs and TPUs. There are several configurations to use this implementation, including standard cluster (where you have full control) and aut

Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1

Model ReleasesDGX agent

arXiv:2607.06764v1 Announce Type: new Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary

Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks

Model ReleasesDGX agent

arXiv:2607.07189v1 Announce Type: new Abstract: Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they can handle the

FedCVESA: Taking Away Training Data in Federated Learning via Correlation Value Encoding and Segmented Aggregation

Model ReleasesDGX agent

arXiv:2607.07314v1 Announce Type: cross Abstract: Federated learning (FL) avoids explicit data exposure by keeping raw data on local clients, yet privacy risks remain in the training process and the l

From Content to Audience: A Multimodal Annotation Framework for Broadcast Television Analytics

Model ReleasesDGX agent

arXiv:2603.26772v2 Announce Type: replace-cross Abstract: Automated semantic annotation of broadcast television content presents distinctive challenges, combining structured audiovisual composition, d

From My View to Yours: Learning Egocentric Cues from Exocentric Video using Privileged Egocentric Supervision

Model ReleasesDGX agent

arXiv:2501.05711v4 Announce Type: replace Abstract: Vision Language Models (VLMs) have achieved strong performance across a wide range of video understanding tasks. However, their viewpoint-invariant

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining

Model ReleasesDGX agent

arXiv:2607.07494v1 Announce Type: cross Abstract: Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining. Communicating gradients in low-precision formats, su

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

Model ReleasesDGX agent

arXiv:2607.06929v1 Announce Type: cross Abstract: Music aesthetic assessment is a challenging yet underexplored problem, requiring models to capture fine-grained, multi-dimensional human perceptual ju

POPS: Recovering Unlearned Multi-Modality Knowledge in MLLMs with Prompt-Optimized Parameter Shaking

Model ReleasesDGX agent

arXiv:2607.06649v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated impressive performance on cross-modal tasks by jointly training on large-scale textual and

Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

Model ReleasesDGX agent

arXiv:2607.07229v1 Announce Type: new Abstract: Prior work has shown that chain-of-thought (CoT) reasoning is often unfaithful: a model's stated reasoning does not reliably reflect the process that pr

Retrieving and Refining Winning Noise Tickets for Diffusion-Based Motion Generation

Model ReleasesDGX agent

arXiv:2607.06843v1 Announce Type: new Abstract: Diffusion-based text-to-motion models synthesize realistic human motions but often exhibit semantic drift from the input text. Motion is inherently temp

8 Jul 2026

ArtisanCAD: An Industrial-Level CAD Agent with Expert-Grounded Knowledge Distillation

Model ReleasesDGX agent

arXiv:2607.05750v1 Announce Type: new Abstract: Computer-aided design (CAD) for industrial components requires long-horizon procedural modeling, robust feature dependencies, editable parametric geomet

Benchmarking KV-Cache Optimizations across Task Quality and System Performance for Long-Context Serving

Model ReleasesDGX agent

arXiv:2607.05399v1 Announce Type: cross Abstract: Large language model serving is increasingly limited by KV-cache growth under long-context workloads, yet existing KV-cache compression techniques are

Exogenous Dropout: A Simple, Strong Baseline for Corruption-Robust Time Series Forecasting with Covariates

Model ReleasesDGX agent

arXiv:2607.05452v1 Announce Type: new Abstract: Time series forecasters that use exogenous covariates are fragile in deployment: when those covariates are noised, temporally misaligned, or missing, st

Generalized Synthetic Image Detection with Enhanced RGB-Noise Representation Learning

Model ReleasesDGX agent

arXiv:2607.06354v1 Announce Type: new Abstract: The rapid advancement of large-scale generative models has accelerated the spread of highly deceptive AI-generated images, making generalized synthetic

Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre-Registered Test

Model ReleasesDGX agent

arXiv:2607.06001v1 Announce Type: new Abstract: We report a pre-registered, two-part experiment on small economies of frontier language-model agents (Claude Opus 4.8), testing two quantitative predict

LLM-Driven Neural Network Generation with Same-Family Architecture Guidance: Disentangling Transfer and Adaptation

Model ReleasesDGX agent

arXiv:2607.05704v1 Announce Type: cross Abstract: Large language models (LLMs) can generate neural-network modifications, but unrestricted generation is often invalid or harmful. This paper studies a

← Previous
1…300301302303304…1036
Next →