AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,332 results
29 Jul 2026

Beyond 'What to Retrieve': Uncertainty in Retrieval-Augmented Code Generation

Model ReleasesDGX agent

arXiv:2607.24884v1 Announce Type: cross Abstract: Repository-level code generation relies on heterogeneous evidence whose relevance, compatibility, and completeness are inherently uncertain. Similar-c

Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization

Model ReleasesDGX agent

arXiv:2607.25451v1 Announce Type: new Abstract: Language models are almost always quantized before they are deployed, and a growing line of work asks whether quantization also lowers their privacy ris

BREAKING: Grok 4.5 (high) ranks #1 on the HighWalk benchmark, which tests how well AI agents update technical specifications from code chang…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

BREAKING: Grok 4.5 (high) ranks #1 on the HighWalk benchmark, which tests how well AI agents update technical specifications from code changes. Grok delivered the best combination of quality and opera

BREAKING: Grok 4.5 just claimed the top spot on the new HighWalk Benchmark. The independent test measures how well AI models update real tec…

Model ReleasesDGX agent

BREAKING: Grok 4.5 just claimed the top spot on the new HighWalk Benchmark. The independent test measures how well AI models update real technical specifications from 46 Laravel commits — heavy on cod

BREAKING: Grok 4.5 ranked #1 on LaurenBench with a score of 56.9%, ahead of Claude Sonnet 5, GLM 5.2, Claude Opus 5, Kimi K3 and GPT-5.6. Th…

Model ReleasesDGX agent

BREAKING: Grok 4.5 ranked #1 on LaurenBench with a score of 56.9%, ahead of Claude Sonnet 5, GLM 5.2, Claude Opus 5, Kimi K3 and GPT-5.6. The benchmark tests real-world AI agents across conversation,

BREAKING: SpaceXAI's newly released Grok Voice Think Fast 2.0 beats voice models from OpenAI, Google, Alibaba, and DeepSlate in the Artifici…

Model ReleasesDGX agent

SpaceXAI has released its new Grok Voice Think Fast 2.0, which on the Artificial Analysis Speech‑to‑Speech benchmark outperformed leading models from OpenAI, Google, Alibaba and DeepSlate. The claim w

Bridging Compute- and Data-Optimal Pretraining

Model ReleasesDGX agent

arXiv:2607.25271v1 Announce Type: cross Abstract: Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in whic

Building Large-Scale English-Romanian Literary Translation Resources with Open Models

Model ReleasesDGX agent

arXiv:2509.07829v4 Announce Type: replace-cross Abstract: Literary translation has recently gained attention as a distinct and complex task in machine translation research, yet translation by small op

Built and released BetterGPT-150M – A compact 150M parameter completion model (+ live HF Space demo)

Model ReleasesDGX agent

Hey everyone, ​I recently finished pre-training BetterGPT-150M, a small, lightweight causal language model with ~152 million parameters.Trained on 15B tokens. Dataset & Training: Trained across stable

CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs

Model ReleasesDGX agent

arXiv:2607.24754v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to provide mental health support, requiring reliable evaluation of safety, empathy, and therapeutic

CD-RMOT-Bench: Benchmarking the Cross-Domain Referring Multi-Object Tracking

Model ReleasesDGX agent

arXiv:2607.25239v1 Announce Type: new Abstract: Referring multi-object tracking (RMOT) extends tracking from category-driven perception to language-guided understanding by grounding object trajectorie

Certified-Gap Dual-Price Policies for Real-Time Truckload Bid Acceptance with Relocating, Clock-Constrained Resources

Model ReleasesDGX agent

arXiv:2607.16891v2 Announce Type: replace-cross Abstract: A truckload carrier must accept or reject each load tender within seconds. The decision depends on fleet state, hours-of-service (HOS) clocks,

Chart-Supported or Model-Supplied? Examining MLLM-Generated Claims for Accessible Visualization

Model ReleasesDGX agent

arXiv:2607.25021v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) can connect visualization patterns to external causes, consequences, and domain knowledge, but the evidential b

ChatGPT Made Me Cry Tonight

Model ReleasesDGX agent

Sorry if flair is wrong. I decided to finally get a ChatGPT subscription after some conversations with it about health issues with my dog. I've only used AI for coding work, primarily Claude, but I fe

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

Model ReleasesDGX agent

arXiv:2607.25294v1 Announce Type: cross Abstract: Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has hig

COCO-OLAC: A Benchmark for Occluded Panoptic Segmentation and Image Understanding

Model ReleasesDGX agent

arXiv:2409.12760v3 Announce Type: replace Abstract: To help address the occlusion problem in panoptic segmentation and image understanding, this paper proposes a new large-scale dataset named COCO-OLA

CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models

Model ReleasesDGX agent

arXiv:2607.24999v1 Announce Type: cross Abstract: LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched

CogEEGAgent: Toward Autonomous Cognitive EEG Analysis with Grounded Execution and Selection-Aware Verification

Model ReleasesDGX agent

arXiv:2607.25045v1 Announce Type: new Abstract: Electroencephalography (EEG) analysis in cognitive studies requires specialized expertise and involves many defensible choices over contrasts, channels,

ComfyUI Face Swap Tutorial: Fast, Clean Results VFX artist @heydoughogan walks through a fully automated face swap workflow built inside Com…

Model ReleasesDGX agent

ComfyUI Face Swap Tutorial: Fast, Clean Results VFX artist @heydoughogan walks through a fully automated face swap workflow built inside ComfyUI - combining Florence 2, SAM2, and WAN Video into a sing

Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference

Model ReleasesDGX agent

arXiv:2607.25018v1 Announce Type: new Abstract: Large language model (LLM) cascades reduce inference cost by routing easy queries to a small model and deferring hard queries to a larger one. Productio

Cooperative Multi-UAV Navigation in Complex Environments via Systematic Multi-Agent Deep Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.25754v1 Announce Type: new Abstract: Cooperative navigation of multi-agent UAVs in complex environments faces key challenges including local optima traps, sparse rewards, learning imbalance

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model

Model ReleasesDGX agent

arXiv:2607.25487v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness b

COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution

Model ReleasesDGX agent

arXiv:2607.25400v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly entrusted with natural-language workflow instructions (e.g., retail-payment policies) that specify no

Crystalis: Progressive Nucleation and Semantic Annealing for Coordinated Multi-View Visualization Generation

Model ReleasesDGX agent

arXiv:2607.24766v1 Announce Type: new Abstract: Large language models (LLMs) can generate individual charts, but coordinated multi-view visualizations (CMVs), where views share data flows and cross-vi

Data Quality Profiling at Scale with Progressive Sampling: A Benchmark for Data-Centric AI Pipelines

Model ReleasesDGX agent

arXiv:2607.25356v1 Announce Type: cross Abstract: Data quality profiling -- computing missing-value rates, duplicate fractions, outlier densities, and functional-dependency violations -- is foundation

Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks

Model ReleasesDGX agent

arXiv:2607.25227v1 Announce Type: cross Abstract: Large Language Models (LLMs) have been widely applied in high-stakes decision-making scenarios such as corporate strategy, and users are increasingly

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

Model ReleasesDGX agent

arXiv:2607.25675v1 Announce Type: new Abstract: Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized a

DeepSeek V4 Flash isn't just for inference anymore. Fine-tune it on Fireworks with supervised fine-tuning, preference tuning, and combined p…

Model ReleasesDGX agent

DeepSeek V4 Flash isn't just for inference anymore. Fine-tune it on Fireworks with supervised fine-tuning, preference tuning, and combined preference optimization from the managed UI. Reinforcement le

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

Model ReleasesDGX agent

arXiv:2607.26041v1 Announce Type: new Abstract: Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success o

Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs

Model ReleasesDGX agent

arXiv:2607.25959v1 Announce Type: cross Abstract: Wikipedia and Wikidata are widely used for information access, LLM pre-training, and retrieval-augmented generation. Their knowledge is deeply connect

Diffusion Model-based Parameter Estimation in Dynamic Power Systems

Model ReleasesDGX agent

arXiv:2411.10431v3 Announce Type: replace Abstract: Parameter estimation, which represents a classical inverse problem, is often ill-posed as different parameter combinations can yield identical outpu

DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization

Model ReleasesDGX agent

arXiv:2607.24856v1 Announce Type: cross Abstract: Social media imagery (SMI) provides timely and fine-grained ground perspectives that are valuable for situational awareness and emergency response. Un

dropped 4k on a spark, am I crazy?

Model ReleasesDGX agent

Saw that the Asus Ascent 1tb was going for $3,950 from a few sources, couldn't stop thinking about it, finally just went ahead and did it. Am I completely insane? Will I regret this? I can't imagine t

Dual-Domain Manifold Modeling for Hyperspectral Image Fusion

Model ReleasesDGX agent

arXiv:2607.25338v1 Announce Type: new Abstract: Achieving a coherent integration of spectral richness and spatial fidelity remains a central objective in hyperspectral image fusion. However, existing

Dual-Level Atomic and Coordination Geometry Learning for Crystal Property Prediction Using Graph Neural Networks

Model ReleasesDGX agent

arXiv:2607.24818v1 Announce Type: cross Abstract: Accurate prediction of crystal properties remains a key challenge in computational materials science. While graph neural networks (GNNs) such as CGCNN

Emergent Latent-State Computation under Stochastic Volatility

Model ReleasesDGX agent

arXiv:2607.25459v1 Announce Type: cross Abstract: Mechanistic interpretability has largely focused on language models and deterministic toy tasks. Much less is known about how sequence models internal

Endpoint Replay: Compressing the Recency Buffer in Deep Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.25123v1 Announce Type: new Abstract: Experience replay remains one of the most practical and useful algorithmic tools in the deep reinforcement learning (DRL) toolbox. Aside from the limite

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

Model ReleasesDGX agent

arXiv:2607.25933v1 Announce Type: cross Abstract: Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practic

Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA

Model ReleasesDGX agent

arXiv:2607.25921v1 Announce Type: cross Abstract: In this work, we study the use of Vision-Language Models (VLMs) for anomaly detection in an agent-driven game Quality Assurance (QA) pipeline focusing

Everyone is unique: Towards Behaviorally Heterogeneous Negotiation Dialogue Systems for Debt Collection

Model ReleasesDGX agent

arXiv:2607.25218v1 Announce Type: new Abstract: Debt collection is a critical negotiation task in the financial industry, with strong practical relevance and exceptional academic value as a behavioral

EviDAG: Auditable Causal DAG Authoring with Biomedical Literature

Model ReleasesDGX agent

arXiv:2607.21859v2 Announce Type: replace Abstract: Constructing causal directed acyclic graphs (DAGs) is a core step in biomedical causal analysis, yet it remains a largely manual process. Analysts m

Fast, accurate, and differentiable: a neural-network surrogate for NRSur7dq4 precessing binary black hole waveforms

Model ReleasesDGX agent

arXiv:2607.24960v1 Announce Type: cross Abstract: We present a neural network surrogate model that emulates the NRSur7dq4 gravitational waveform model for precessing binary black hole mergers. The sur

Few-Shot Open-Vocabulary Remote Sensing Segmentation via Textual Inversion

Model ReleasesDGX agent

arXiv:2607.25563v1 Announce Type: new Abstract: Open-vocabulary segmentation labels arbitrary categories from a text query without per-class training, yet on remote sensing imagery it underperforms on

First Kimi K3 results on home lab ~ 4t/s

Model ReleasesDGX agent

I've got better results than expected for 768gb DDR5 and 2x5090. Using fork https://github.com/pwilkin/llama.cpp/tree/kimi-k3-text and https://huggingface.co/GrEarl/Kimi-K3-GGUF Q2_K quant. Prefill sp

Food Image Segmentation with LLM-Derived Ingredient Labels and Multimodal Fusion

Model ReleasesDGX agent

arXiv:2607.25820v1 Announce Type: new Abstract: Food image segmentation plays a vital role in health-related applications such as nutrition tracking and personalized health monitoring. However, existi

Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact

Model ReleasesDGX agent

arXiv:2607.25589v1 Announce Type: cross Abstract: Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and reposito

From Cellular Responses to Pharmacological Domains: Multimodal Zero-Shot Drug Representation Learning

Model ReleasesDGX agent

arXiv:2607.25322v1 Announce Type: new Abstract: Multimodal drug discovery enables drug representation learning beyond chemical structure by incorporating cellular responses such as gene expression and

From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

Model ReleasesDGX agent

Figure 1: CUDA-to-MLX optimization translation map. CUDA optimization knowledge can be translated into architecture-native MLX strategies rather than copied instruction-for-instruction. We face a new

From Deterministic to Generative Deep Learning for Urban Air Quality Reconstruction from Sparse Observations

Model ReleasesDGX agent

arXiv:2607.25687v1 Announce Type: cross Abstract: Full-field reconstruction of air pollution is essential for evaluating pollution exposure and supporting public health decision-making. However, the c

GAUGE: Grading Agent-Built Financial Models Without a Golden Answer

Model ReleasesDGX agent

arXiv:2607.24889v1 Announce Type: cross Abstract: Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be checked mechani

Generalization from Low- to Moderate-Resolution Spectra with Neural Networks for Stellar Parameter Estimation: A Case Study with DESI

Model ReleasesDGX agent

arXiv:2602.15021v2 Announce Type: replace-cross Abstract: Cross-survey generalization is a critical challenge in stellar spectral analysis, particularly in cases such as transferring from low- to mode

GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games? We…

Model ReleasesDGX agent

GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games? We investigated. The harness was not letting it remember what

GraphRareBench: An Auditable Graph-Evidence Benchmark for Phenotype-Driven Rare-Disease Diagnosis

Model ReleasesDGX agent

arXiv:2607.24878v1 Announce Type: cross Abstract: Phenotype-driven diagnostic benchmarks usually report the rank of the reference disease, but they rarely reveal which plausible alternatives are ranke

GrocLM: Grocery Category Recommendation in E-Commerce with Large Language Models

Model ReleasesDGX agent

arXiv:2607.24764v1 Announce Type: new Abstract: The rapid growth of online grocery shopping requires recommendation systems that capture cyclical purchasing behavior and diverse user intents. Traditio

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Model ReleasesDGX agent

arXiv:2607.25398v1 Announce Type: new Abstract: Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context,

Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed

Model ReleasesDGX agent

arXiv:2607.24898v1 Announce Type: cross Abstract: State-of-the-art toxicity detectors for text-to-image generation adopt a one-size-fits-all approach: a single universal model applying fixed safety gu

HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising

Model ReleasesDGX agent

arXiv:2607.24779v1 Announce Type: new Abstract: Online advertising bidding systems typically deploy multiple offline-trained expert models (e.g., PID controllers, model predictive control, offline RL

How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model

Model ReleasesDGX agent

arXiv:2607.25583v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) and low-bit quantization are now standard tools for adapting language models under tight compute budgets, yet the

https://x.com/EMostaque/status/2082600218529235174

Model ReleasesDGX agent

After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run. The results: - 20% lower serving costs from production GPU kernel improvements. -

I pre-trained a 700m on 18B tokens optimized for Python and Wikitext | TheOneWhoWill/Shibai-700M-Base · Hugging Face

Model ReleasesDGX agent

I know this is the 1000000th new sub billion parameter model out there and probably isn't as good as Qwen 3 0.6B or Qwen 3.5 0.8B but it still packs a decent punch. My intention to to continuously pre

← Previous
1…5253545556…373
Next →