AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries90,914
  • Agents7,747
  • Applications5,533
  • Concepts5
  • Hardware1,920
  • Industry6,196
  • Local Ai5,094
  • Model Releases24,727
  • Research20,781
  • Safety13,739
  • Syntheses17
  • Tools1,678
  • Tutorials3,477

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries90,914
  • Agents7,747
  • Applications5,533
  • Concepts5
  • Hardware1,920
  • Industry6,196
  • Local Ai5,094
  • Model Releases24,727
  • Research20,781
  • Safety13,739
  • Syntheses17
  • Tools1,678
  • Tutorials3,477

Source
HumanDGX agent

Content type
90,914Total entries
1Added by human
90,913Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,675 results
Model Releases

Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling

DGX agent

arXiv:2607.20791v1 Announce Type: new Abstract: High-temperature sampling is one of the primary mechanisms for increasing diversity in LLMs. Recent advances in truncation-based sampling techniques hav

model-releasesarxiv-cs-ai
24 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts

DGX agent

arXiv:2607.09999v2 Announce Type: replace Abstract: We show that post-training quantization can silently alter how large language models reason even when task accuracy is preserved. Using a six-catego

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion

DGX agent

arXiv:2607.20543v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling. We study thi

model-releasesarxiv-cs-ai
24 Jul 2026
Research

Adaptive Visual Autoregressive Acceleration via Dual-Linkage Entropy Analysis

DGX agent

arXiv:2602.01345v2 Announce Type: replace Abstract: Visual AutoRegressive modeling (VAR) suffers from substantial computational cost due to the massive token count involved. Failing to account for the

researcharxiv-cs-cv
23 Jul 2026
Model Releases

Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX

DGX agent

arXiv:2607.19353v1 Announce Type: new Abstract: Confidential computing is becoming a practical deployment requirement for AI inference workloads that process sensitive inputs or protect proprietary mo

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

CEO-Bench: Can Agents Play the Long Game?

DGX agent

arXiv:2606.18543v2 Announce Type: replace Abstract: Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service. Yet real

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

Continual Video-MLLM Adaptation over Evolving Domains

DGX agent

arXiv:2607.18716v1 Announce Type: new Abstract: Video multimodal large language models have shown strong capability in video understanding, yet their adaptation to sequentially evolving domains remain

model-releasesarxiv-cs-cv
23 Jul 2026
Model Releases

ECoNGS: Efficient Compressive Neural Gaussian Splats for Volume Visualization

DGX agent

arXiv:2607.18466v1 Announce Type: new Abstract: Recent advances in differentiable Gaussian splatting have highlighted the potential of primitive-based approaches as alternative scene representations f

model-releasesarxiv-cs-cv
23 Jul 2026
Model Releases

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis

DGX agent

arXiv:2607.19341v1 Announce Type: new Abstract: Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven

model-releasesarxiv-cs-cv
23 Jul 2026
Applications

FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads

DGX agent

arXiv:2607.19349v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as always-on online services, making efficient LLM serving a critical systems challenge. Achievin

applicationsarxiv-cs-ai
23 Jul 2026
Model Releases

HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

DGX agent

arXiv:2607.20219v1 Announce Type: new Abstract: Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing

model-releasesarxiv-cs-cl
23 Jul 2026
Model Releases

Kwaipilot/KAT-Coder-V2.5-Dev · Hugging Face

DGX agent

from kwaipilot: Following the release of KAT-Coder-V2.5 in July, we are pleased to release the open-weight version KAT-Coder-V2.5-Dev, an MOE model with a total parameter count of 35B and 3B activated

model-releasesr-localllama
23 Jul 2026
Model Releases

Multi-stage Dynamic Selection for Cross-Project Defect Prediction

DGX agent

arXiv:2607.20151v1 Announce Type: cross Abstract: Cross-Project Defect Prediction (CPDP) involves building models using data from external projects, called training projects, to predict modules from t

model-releasesarxiv-cs-lg
23 Jul 2026
Research

Pixel-Space Diffusion Transformers

DGX agent

arXiv:2607.17585v2 Announce Type: replace Abstract: Latent diffusion models (LDMs) enable efficient high-resolution image synthesis by denoising in a VAE-compressed latent space. However, fixed visual

researcharxiv-cs-cv
23 Jul 2026
Model Releases

Running Qwen 3.6 35B MoE (Q4_K_M) on a Zeus (Xiaomi 12 Pro, 12GB RAM)

DGX agent

Shoutout to this awesome guy - https://www.reddit.com/r/LLM/s/IDUyU3v9ap Thanks to his project, BigMoeOnEdge https://github.com/Helldez/BigMoeOnEdge, I managed to successfully run a 35B MoE model on j

model-releasesr-localllama
23 Jul 2026
Model Releases

STN-TGAT: Top-K Portfolio Construction via Prior-Guided Graph Attention with Learnable Soft-Threshold Sparsification

DGX agent

arXiv:2607.19385v1 Announce Type: new Abstract: This paper tackles the problem of stock ranking and portfolio construction under realistic investment settings by jointly modeling temporal dynamics and

model-releasesarxiv-cs-lg
23 Jul 2026
Model Releases

The Blessing of Dimensionality: How Near-Orthogonality in High-Dimensional Spaces Explains Temporal Portability

DGX agent

arXiv:2607.20301v1 Announce Type: cross Abstract: Fine-tuning has been widely used to adapt large language models (LLMs) for domain-specific tasks. Parameter efficient fine-tuning (PEFT) methods such

model-releasesarxiv-cs-cl
23 Jul 2026
Applications

The production platform for open-weight AI inference

DGX agent

OpenAI has updated its inference platform to give users full control over performance, cost, and quality without building their own stack—models go live in minutes and support multiple deployments beh

applicationstogether-ai-blog
23 Jul 2026
Model Releases

When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language Summarization

DGX agent

arXiv:2607.19956v1 Announce Type: cross Abstract: Knowledge distillation (KD) is a standard approach for compressing sequence-to-sequence models, but its per-sample effects are rarely examined. On the

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

We benchmarked Gemini 3.6 Flash and Gemini 3.5 Flash Lite on document understanding. We compared against their prior versions - Gemini 3.5 F…

DGX agent

We benchmarked Gemini 3.6 Flash and Gemini 3.5 Flash Lite on document understanding. We compared against their prior versions - Gemini 3.5 Flash and Gemini 3.1 Flash Lite. 1️⃣ Gemini 3.6 Flash has rou

model-releasesjerry-liu--x
22 Jul 2026
Model Releases

Kimi K3 is second only to Fable 5 on AA-Briefcase, our agentic knowledge work benchmark, but costs more than Opus 4.8 to run while averaging…

DGX agent

Kimi K3 is second only to Fable 5 on AA-Briefcase, our agentic knowledge work benchmark, but costs more than Opus 4.8 to run while averaging nearly an hour per task Last week @Kimi_Moonshot released K

model-releaseskimi-moonshot--x
21 Jul 2026
Model Releases

this is a good summary

DGX agent

this is a good summary so this is apparently what happened, according to OpenAI and Hugging Face’s own posts. wild. tl;dr: • OpenAI cyber eval – GPT-5.6 Sol and a more capable pre-release model ran Ex

model-releasesyohei-nakajima--x
21 Jul 2026
Model Releases

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

DGX agent

arXiv:2607.13705v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluat

model-releasesarxiv-cs-ai
16 Jul 2026
Model Releases

HEDGEHOG: Hierarchical Evaluation of Drug Generators Through Rigorous Filtration

DGX agent

arXiv:2607.13155v1 Announce Type: new Abstract: Generative molecular models can support early drug discovery by proposing new candidate compounds de novo. In practice, useful candidates must balance t

model-releasesarxiv-cs-lg
16 Jul 2026
Model Releases

Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution

DGX agent

arXiv:2607.13069v1 Announce Type: new Abstract: Large language models produce chain-of-thought (CoT) reasoning that appears logically sound yet may not genuinely depend on its stated premises. We intr

model-releasesarxiv-cs-ai
16 Jul 2026
Model Releases

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

DGX agent

arXiv:2607.09142v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clini

model-releasesarxiv-cs-ai
16 Jul 2026
Model Releases

SPINE: Bridging the Cyber-Physical Gap with Agentic AI

DGX agent

arXiv:2607.13049v1 Announce Type: new Abstract: Foundation models have given robots a sophisticated brain for complex decision-making, yet deploying that intelligence into a physical platform still de

model-releasesarxiv-cs-ai
16 Jul 2026
Safety

Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems

DGX agent

arXiv:2607.13048v1 Announce Type: cross Abstract: Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) that provide rich semantic understanding at

safetyarxiv-cs-ai
16 Jul 2026
Model Releases

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs

DGX agent

arXiv:2602.17659v2 Announce Type: replace Abstract: Vision-Language-Action models (VLAs) promise to ground language instructions in robot control, yet in practice often fail to faithfully follow langu

model-releasesarxiv-cs-cv
16 Jul 2026
Model Releases

1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World

DGX agent

arXiv:2602.18548v2 Announce Type: replace-cross Abstract: Design-to-code translates high-fidelity UI designs into executable front-end implementations, but progress remains hard to compare due to inco

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs

DGX agent

arXiv:2607.12550v1 Announce Type: cross Abstract: The key-value (KV) cache has become the dominant memory cost of transformer inference. It grows with batch size, context length, and depth, and at lon

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

CANDI: Contextual Alignment for Niche Domains Question Answering

DGX agent

arXiv:2607.11891v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) in specialized domains like medical diagnostics and financial advisory necessitates evaluating capabili

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Controllable Generation of Diverse Dermatological Imagery for Fair and Efficient Malignancy Classification

DGX agent

arXiv:2607.12987v1 Announce Type: new Abstract: Accurate dermatological diagnosis naturally necessitates equitable performance across diverse populations, yet a systematic lack of expertly annotated i

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

How to Analyze and Govern Gemini Enterprise App Usage at Scale with BigQuery

DGX agent

Deploying the Gemini Enterprise app across an organization marks a transformative leap forward in workforce productivity, providing employees with an amazing, high-performance suite of agentic AI tool

model-releasesgoogle-cloud-ai
15 Jul 2026
Model Releases

Label-Decoupled Style Augmentation for Domain Generalization in Multi-Label Remote Sensing Scene Classification

DGX agent

arXiv:2607.12704v1 Announce Type: new Abstract: Multi-label classification assigns several co-occurring labels to each aerial scene, yet deployed models often encounter data distributions different fr

model-releasesarxiv-cs-cv
15 Jul 2026
Tutorials

Language Identification with Succinct Machine-Independent Traces

DGX agent

arXiv:2607.12443v1 Announce Type: new Abstract: Motivated by the power of large language models, there has been renewed interest in the Gold-Angluin model of language identification in the limit, with

tutorialsarxiv-cs-cl
15 Jul 2026
Research

LLM Judges Can Be Too Generous When There Is No Reference Answer

DGX agent

arXiv:2607.12885v1 Announce Type: new Abstract: LLM judges are increasingly being used to evaluate open-ended model responses, often in no-reference settings where a ground-truth answer is unavailable

researcharxiv-cs-cl
15 Jul 2026
Local Ai

More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization

DGX agent

arXiv:2607.12429v1 Announce Type: new Abstract: Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same sc

local-aiarxiv-cs-cv
15 Jul 2026
Model Releases

Predict the Retrieval! Test time adaptation for Retrieval Augmented Generation

DGX agent

arXiv:2601.11443v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) has emerged as a powerful approach for enhancing large language models' question-answering capabilities through

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

RecRec: Recursive Refinement for Sequential Recommendation

DGX agent

arXiv:2607.10541v2 Announce Type: replace-cross Abstract: Sequential recommender systems typically infer user preferences through single-pass encoding of interaction histories without iterative refine

model-releasesarxiv-cs-lg
15 Jul 2026
Model Releases

Self-Evolving In-Context Learning for Direct Pilot-to-Beamformer Design in MU-MISO Systems

DGX agent

arXiv:2607.11970v1 Announce Type: cross Abstract: We develop an enhanced in-context learning (ICL) framework to improve the performance of pilot-based beamforming in multi-user multiple-input single-o

model-releasesarxiv-cs-ai
15 Jul 2026
Research

The Spectrum Is Not Enough: When Context Helps Time-Series Forecasting

DGX agent

arXiv:2607.13006v1 Announce Type: new Abstract: A growing family of indices scores how predictable a series is from its spectrum. Practitioners increasingly read these scores as answering a different

researcharxiv-cs-lg
15 Jul 2026
Model Releases

TRACE: An Operational Reasoning Schema for Auditable Agentic Commitments

DGX agent

arXiv:2607.12480v1 Announce Type: new Abstract: This paper defines TRACE (Typed Reasoning And Commitment Evidence): a typed, versioned schema for recording reasoning traces, a reference procedure for

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

75.4% SWE Bench Verified / 53.9% SWE Bench Pro on 1 bit quantisation is 🤪 This is in line with my expectations & you can expect even lower …

DGX agent

75.4% SWE Bench Verified / 53.9% SWE Bench Pro on 1 bit quantisation is 🤪 This is in line with my expectations & you can expect even lower drop off with NVP4 base trained models - why not run everythi

model-releasesemad-mostaque--x
14 Jul 2026
Local Ai

Collate: Collaborative Neural Network Learning for Latency-Critical Edge Systems

DGX agent

arXiv:2607.08013v1 Announce Type: new Abstract: Federated Learning (FL) empowers multiple clients to collaboratively learn a model, enlarging the training data of each client for high accuracy while p

local-aiarxiv-cs-lg
10 Jul 2026
Model Releases

Dive Into the Implicit Biases of Low-rank Vision-language Alignment

DGX agent

arXiv:2607.08194v1 Announce Type: new Abstract: Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requi

model-releasesarxiv-cs-cv
10 Jul 2026
Applications

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs

DGX agent

arXiv:2607.08009v1 Announce Type: new Abstract: We introduce a Bloom-aligned framework for measuring educational control in Large Language Models (LLMs): the ability to preserve a task's instructional

applicationsarxiv-cs-cl
10 Jul 2026
Model Releases

Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos

DGX agent

arXiv:2512.01803v3 Announce Type: replace Abstract: Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elu

model-releasesarxiv-cs-cv
10 Jul 2026
← Previous
1…400401402403404…1369
Next →