AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,555 results
Model Releases

Deep4ge: DNN Training Trajectories for Fault Detection and Diagnosis

DGX agent

arXiv:2607.12868v1 Announce Type: cross Abstract: Deep learning systems often fail due to subtle implementation faults that alter training behavior. Recent work has studied how to detect and diagnose

model-releasesarxiv-cs-lg
15 Jul 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents

DGX agent

arXiv:2509.21842v2 Announce Type: replace Abstract: Travel planning (TP) agent has recently worked as an emerging building block to interact with external tools/resources for travel itinerary generati

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

DeGuNet: Depth-Guided Ultra-Compact Backbones for Efficient LiDAR-Camera 3D Detection

DGX agent

arXiv:2607.12419v1 Announce Type: new Abstract: In autonomous driving perception, the fusion of LiDAR and camera modalities has become the dominant paradigm for 3D object detection. However, current m

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability

DGX agent

arXiv:2607.12056v1 Announce Type: new Abstract: Online shopping is increasingly shifting toward a model in which AI agents independently search for products, compare options, evaluate constraints, and

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection

DGX agent

arXiv:2607.11969v1 Announce Type: cross Abstract: Point-adjustment (PA), long the default scoring protocol in time-series anomaly detection (TSAD), was shown by Kim et al. (2022) to award near-perfect

model-releasesarxiv-cs-lg
15 Jul 2026
Model Releases

DiffRadar: Differentiable Physics-Aware Radar SLAM with Gaussian Fields

DGX agent

arXiv:2607.12265v1 Announce Type: new Abstract: Radar sensing is increasingly used in mobile systems because it operates reliably under poor lighting, adverse weather, and privacy-sensitive settings w

model-releasesarxiv-cs-ro
15 Jul 2026
Model Releases

DiTailed: Ensuring Visual Object Consistency in Text-Image-to-Image Flow Matching Models

DGX agent

arXiv:2607.12539v1 Announce Type: new Abstract: Despite remarkable progress in text-guided image editing, generative models frequently fail to preserve visual object consistency, defined as the preser

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution

DGX agent

arXiv:2607.13034v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task act

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory

DGX agent

arXiv:2603.25112v2 Announce Type: replace-cross Abstract: Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accuracy)

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?

DGX agent

arXiv:2607.12787v1 Announce Type: new Abstract: Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enab

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Dynamic Resource Allocation for Ensemble Determinization MCTS

DGX agent

arXiv:2607.13007v1 Announce Type: new Abstract: Simulation-based algorithms are especially suited for high-uncertainty environments such as adversarial board games with significant elements of randomn

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Edge-Aware Thermal Infrared UAV Swarm Tracking

DGX agent

arXiv:2607.12544v1 Announce Type: new Abstract: Thermal infrared (TIR) imaging is essential for UAV swarm operations in visually degraded environments. However, tracking tiny UAVs remains challenging

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

Efficiently Learning Branching Networks for Multitask Algorithmic Reasoning

DGX agent

arXiv:2512.01113v2 Announce Type: replace-cross Abstract: Algorithmic reasoning -- the ability to perform step-by-step logical inference -- is a synthetic benchmark for evaluating multi-step reasoning

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Egocentric Bias in Vision-Language Models

DGX agent

arXiv:2602.15892v2 Announce Type: replace-cross Abstract: Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce FlipSet

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Enabling 24-hour Agricultural Robotics: Unsupervised Day-to-Night Cross-Modal Image Translation for Nighttime Visual Navigation

DGX agent

arXiv:2607.12065v1 Announce Type: cross Abstract: While visual navigation has been extensively studied in agricultural robotics, most existing systems assume daytime conditions. In fact, deploying aut

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Environment Parameter Gradient Theorem for Policy-Environment Co-Design in Reinforcement Learning

DGX agent

arXiv:2607.12590v1 Announce Type: cross Abstract: Reinforcement learning (RL) is traditionally concerned with learning a control policy for a fixed environment. In many engineering systems, however, t

model-releasesarxiv-cs-lg
15 Jul 2026
Model Releases

Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models

DGX agent

arXiv:2607.12739v1 Announce Type: new Abstract: A language model may be asked either what experts believe about a contested claim or what it believes about the claim itself. A trustworthy conversation

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

Evaluating Health Misinformation in Low-Resource Languages: Integrating Small Language Models with a Culturally-Sensitive Responsible NLP Framework (Bangla as a Case Study)

DGX agent

arXiv:2607.12336v1 Announce Type: cross Abstract: Artificial Intelligence (AI) technologies, while serving as a foundational enabler for modern social media and digital health services, exert a bivale

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations

DGX agent

arXiv:2607.12884v1 Announce Type: new Abstract: Patients seeking medical information often ask questions that embed incorrect assumptions or misconceptions. In such cases, safe medical communication r

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

Evidence-Grounded AI for Musculoskeletal Care

DGX agent

arXiv:2607.12527v1 Announce Type: new Abstract: Musculoskeletal diseases are among the leading causes of disability worldwide and create the greatest global need for rehabilitation. Because recovery,

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

DGX agent

arXiv:2509.24372v3 Announce Type: replace-cross Abstract: Fine-tuning large language models (LLMs) for downstream tasks is an essential stage of modern AI deployment. Reinforcement learning (RL) has e

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction

DGX agent

arXiv:2607.12584v1 Announce Type: cross Abstract: The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics. While rec

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Exploring Zero-Shot Foundation Models for Multivariate Time Series Anomaly Detection

DGX agent

arXiv:2607.12454v1 Announce Type: new Abstract: Multivariate Time Series Anomaly Detection (MTSAD) is essential for reliability and safety in domains such as industrial process monitoring and financia

model-releasesarxiv-cs-lg
15 Jul 2026
Model Releases

Extractable Memorization From First Principles

DGX agent

arXiv:2607.12649v1 Announce Type: cross Abstract: Recent work on extractable memorization in LLMs suffers from two contrasting validity problems. Some studies overstate extraction, e.g., relying on se

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

FairCoder: Probing LLM Bias in High-Stakes Decision Making via Coding Tasks

DGX agent

arXiv:2501.05396v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used in high-stakes decisions such as hiring and college admissions, making their social bias a critic

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality

DGX agent

arXiv:2607.12252v1 Announce Type: new Abstract: Deep research agents are increasingly used to produce long-form financial reports, yet large-scale evaluation remains bottlenecked by the need for human

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

DGX agent

arXiv:2512.01241v4 Announce Type: replace-cross Abstract: Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Forgetful Attention: A Trainable Support-Vector Memory with Certified Selection and Exact Unlearning

DGX agent

arXiv:2607.12204v1 Announce Type: new Abstract: Attention can be viewed as an online learner over context, yet existing test-time memories cannot certify that dropping a token leaves outputs unchanged

model-releasesarxiv-cs-lg
15 Jul 2026
Model Releases

FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem Generation

DGX agent

arXiv:2607.12982v1 Announce Type: new Abstract: Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large Language Models (MLLMs), however analytic geometry remai

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation

DGX agent

arXiv:2607.12687v1 Announce Type: cross Abstract: LLMs can perform language-based quantitative prediction from unstructured inputs, but remain susceptible to hallucinations and overconfident errors, m

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

From Many to Meaningful: Feature-Guided Zero-Shot Chronic Kidney Disease Screening Using Large Language Models

DGX agent

arXiv:2607.12260v1 Announce Type: new Abstract: Early screening of chronic kidney disease (CKD) is essential for preventing irreversible progression; however, many machine learning (ML)-based screenin

model-releasesarxiv-cs-lg
15 Jul 2026
Model Releases

GAINS: Gaussian-based Inverse Rendering from Sparse Multi-View Captures

DGX agent

arXiv:2512.09925v2 Announce Type: replace Abstract: Recent advances in Gaussian Splatting-based inverse rendering extend Gaussian primitives with shading parameters and physically grounded light trans

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents

DGX agent

arXiv:2607.11192v2 Announce Type: replace Abstract: A large share of day-to-day work in professional domains happens inside PDF files: benefits packets, leases, datasheets, clinical guidelines, constr

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

Generating Developable 3D Molecules via Pocket-Conditioned Diffusion and Property-Aware Optimization

DGX agent

arXiv:2607.12349v1 Announce Type: new Abstract: Drug discovery and development is time-consuming and resource-intensive, motivating computational approaches such as diffusion models for de novo drug d

model-releasesarxiv-cs-lg
15 Jul 2026
Model Releases

Generic AI models are not built for medicine. Doximity Ask, trained and served on Fireworks, outperformed GPT-5.6 Sol, Claude Fable 5, and O…

DGX agent

Generic AI models are not built for medicine. Doximity Ask, trained and served on Fireworks, outperformed GPT-5.6 Sol, Claude Fable 5, and OpenEvidence in a Stanford-Harvard clinical AI safety study.

model-releasesfireworks-ai--x
15 Jul 2026
Model Releases

ggml-zendnn : add Q8_0 quantization support by z-sachin · Pull Request #23414 · ggml-org/llama.cpp

DGX agent

Benchmark Results Benchmark configuration: threads = 96 type_k = bf16 type_v = bf16 Llama-3.1-8B-Instruct Q8_0 Prompt Size GGML_CPU_Q8_0 t/s ZenDNN_Q8_0 t/s Gain 256 472.28 730.87 54.75% 512 450.86 83

model-releasesr-localllama
15 Jul 2026
Model Releases

Given GB300 I would estimate this is was trained on about 1e25 flops (same as DeepSeek v4) over 1.6m hours (1 month on 2k chips/28 racks) Co…

DGX agent

Given GB300 I would estimate this is was trained on about 1e25 flops (same as DeepSeek v4) over 1.6m hours (1 month on 2k chips/28 racks) Cost 10-20m ($6-12/hour/chip) The lite version likely 4x less,

model-releasesemad-mostaque--x
15 Jul 2026
Model Releases

Good reason to try Grok 4.5 with Grok Build. It gets better every day!

DGX agent

Good reason to try Grok 4.5 with Grok Build. It gets better every day! Grok 4.5 just took the #1 spot on the Long-Horizon Terminal-Bench, outperforming Claude Fable 5, Claude Opus 4.8 and GPT-5.6-sol

model-releaseselon-musk--x
15 Jul 2026
Model Releases

Graph-Constrained Policy Learning for Extreme Clinical Code Prediction

DGX agent

arXiv:2607.11954v1 Announce Type: cross Abstract: Clinical code prediction maps unstructured discharge summaries to ICD-10-CM leaf codes in a large, sparse, and deeply hierarchical label space. Most s

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

GRID: Grammar-Railed Decoding for Enterprise SQL Generation

DGX agent

arXiv:2607.11951v1 Announce Type: new Abstract: Large language models can write SQL, but enterprise deployment demands more than plausible text: outputs must be syntactically valid, must respect per-r

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Grok 4.5 is worth trying

DGX agent

Grok 4.5 is worth trying Grok 4.5 just ranked #2 on the FrontierSWE leaderboard, outperforming Claude Opus 4.8 and GPT-5.5 It delivers top-tier coding performance with exceptional speed, strong token

model-releaseselon-musk--x
15 Jul 2026
Model Releases

Hermes on Android (Graphene OS)

DGX agent

https://youtu.be/oxpGq5FITgA?si=nkHWLReGCDYe7QfL I got Hermes running in the native Debian Terminal in Graphene OS and its really slick. Voice dictation works amazingly. Im using a remote Hermes gatew

model-releasesr-localllama
15 Jul 2026
Model Releases

How I tricked Claude into leaking your deepest, darkest secrets

DGX agent

How I tricked Claude into leaking your deepest, darkest secrets I've been impressed by the way the Claude web_fetch tool is designed to avoid data exfiltration attacks. Ayush Paul found a hole in that

model-releasessimon-willison
15 Jul 2026
Model Releases

How Inference Compute Shapes Frontier LLM Evaluation

DGX agent

arXiv:2606.17930v2 Announce Type: replace Abstract: AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result,

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks

DGX agent

arXiv:2607.12338v1 Announce Type: new Abstract: Agent benchmarks often compare two agents after all tasks have run, but costly evaluations make partial runs tempting. A task fraction alone does not sh

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

How to Analyze and Govern Gemini Enterprise App Usage at Scale with BigQuery

DGX agent

Deploying the Gemini Enterprise app across an organization marks a transformative leap forward in workforce productivity, providing employees with an amazing, high-performance suite of agentic AI tool

model-releasesgoogle-cloud-ai
15 Jul 2026
Model Releases

How to Realize Recursively Self-Improving Agents and Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture

DGX agent

arXiv:2607.12254v1 Announce Type: new Abstract: Large language model (LLM) agents can increasingly plan, use tools, maintain memory, and execute long-horizon tasks. These advances motivate two linked

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

How to solve PostgreSQL multilingual full-text search limitations with AlloyDB AI

DGX agent

AlloyDB powers enterprise-grade search for some of the largest organizations, providing robust hybrid search capabilities that combine text, vector, and keyword searches into a simple ranked SQL query

model-releasesgoogle-cloud-ai
15 Jul 2026
← Previous
1…102103104105106…470
Next →