AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,499 results
Model Releases

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

DGX agent

arXiv:2607.27084v1 Announce Type: new Abstract: Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in

model-releasesarxiv-cs-cv
30 Jul 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

DGX agent

arXiv:2607.26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

DGX agent

arXiv:2607.26326v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. How

model-releasesarxiv-cs-cv
30 Jul 2026
Model Releases

Semantic-Aware Temporal Adaptation for UAV Anti-UAV Tracking

DGX agent

arXiv:2607.26511v1 Announce Type: new Abstract: UAV Anti-UAV tracking is an emerging low-altitude security task for localizing an adversarial UAV using the onboard camera of a moving observer UAV. It

model-releasesarxiv-cs-cv
30 Jul 2026
Model Releases

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning

DGX agent

arXiv:2607.26873v1 Announce Type: new Abstract: Test-time reinforcement learning (TTRL) enables language models to self-evolve at inference time without labeled feedback. Existing methods rely on answ

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

DGX agent

arXiv:2607.27056v1 Announce Type: cross Abstract: Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retriev

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

DGX agent

arXiv:2607.26173v1 Announce Type: new Abstract: Alignment training, model organisms, and toy models are usually treated as separate research areas. But projects in all three frequently use supervised

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

Simultaneous Coverage and Efficiency Guarantee in Online Conformal Prediction

DGX agent

arXiv:2607.26577v1 Announce Type: new Abstract: Adaptive conformal inference (ACI) of Gibbs and Cand{es and its variants are the standard approach to online conformal prediction under distribution shi

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

SpatialQ: Understanding 3D Gaussian Splatting Scene Quality via Visual-based MLLM

DGX agent

arXiv:2607.26595v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has emerged as an effective representation for novel view synthesis and 3D scene reconstruction, creating an increasing dem

model-releasesarxiv-cs-cv
30 Jul 2026
Model Releases

Structurally Separated Uncertainty in Supervised Latent Variable Models

DGX agent

arXiv:2602.11219v2 Announce Type: replace Abstract: Predictive uncertainty is commonly decomposed into epistemic and aleatoric components, but standard decompositions often produce strongly correlated

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

sure we lose money on every inference but we make it up in volume

DGX agent

sure we lose money on every inference but we make it up in volume We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices f

model-releasesgary-marcus--x
30 Jul 2026
Model Releases

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

DGX agent

arXiv:2607.26355v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly embedded in everyday life and widely used for information seeking, raising concerns about their potential

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method

DGX agent

arXiv:2607.26924v1 Announce Type: new Abstract: Recent work on LeWorldModel (LeWM) has shown that the Sketched Isotropic Gaussian Regularizer (SIGReg) enables stable end-to-end world-model learning fr

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

The Art of Not Forgetting A Local Learning Architecture for Continual Learning

DGX agent

arXiv:2607.26523v1 Announce Type: new Abstract: We introduce CMP (Cognitive Memory Primitive), a continual-learning architecture that repre?sents inputs as sparse relational codes, stores them in a tw

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness

DGX agent

arXiv:2503.10647v2 Announce Type: replace Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dim

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Thinking Machines just released Inkling-Small: 276B total, 12B active. A faster Inkling that matches or beats its 975B big brother in many b…

DGX agent

Thinking Machines just released Inkling-Small: 276B total, 12B active. A faster Inkling that matches or beats its 975B big brother in many benchmarks. To test its speed, we plugged it into HF's speech

model-releasessoumith-chintala--x
30 Jul 2026
Model Releases

This is absolutely wild... Anthropic reviewed their logs and found out that their own supposedly-sandboxed cyber evals had hacked three sepa…

DGX agent

This is absolutely wild... Anthropic reviewed their logs and found out that their own supposedly-sandboxed cyber evals had hacked three separate companies back in April without them noticing! In a rev

model-releasessimon-willison--x
30 Jul 2026
Model Releases

Tight Generalization Bound for AdaBoost

DGX agent

arXiv:2607.26838v1 Announce Type: new Abstract: In this paper we show that the generalization error of AdaBoost is Thetaig(frac{dln(ngamma^{2}/d)}{ngamma^2}+frac{ln(1/elta)}{n}ig), where gamma is the

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans

DGX agent

arXiv:2603.09971v2 Announce Type: replace Abstract: We present TiPToP, a modular manipulation system that integrates pretrained foundation models with a GPU-accelerated Task and Motion Planner to solv

model-releasesarxiv-cs-ro
30 Jul 2026
Model Releases

Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels

DGX agent

arXiv:2607.26121v1 Announce Type: new Abstract: Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Because failures

model-releasesarxiv-cs-ro
30 Jul 2026
Model Releases

ToxScreen: Detecting Whether an LLM Has Been Poisoned

DGX agent

arXiv:2607.26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that c

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

TPCD: Tone-Pressure Contrastive Decoding and the Label-Free Gating Bottleneck in Vision-Language Models

DGX agent

arXiv:2607.26536v1 Announce Type: new Abstract: High-pressure prompts can push vision-language models (VLMs) into unsupported commitments, such as reading illegible text, reporting indeterminate times

model-releasesarxiv-cs-cv
30 Jul 2026
Model Releases

Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation

DGX agent

arXiv:2510.00192v3 Announce Type: replace Abstract: Low-rank adaptation (LoRA) has become a widely used paradigm for parameter-efficient fine-tuning of large language models, yet its representational

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

Transformers Can Learn Rules They've Never Seen: Proof of Computation Beyond Interpolation

DGX agent

arXiv:2603.17019v2 Announce Type: replace Abstract: A central question in the debate over large language models is whether transformers can learn rules they have never seen, or whether they can only i

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

TreeCCA: Canonical Correlation Analysis via Gradient-Boosted Trees

DGX agent

arXiv:2607.27027v1 Announce Type: new Abstract: Gradient-boosted trees dominate tabular machine learning, yet canonical correlation analysis has always relied on linear or neural encoders. We propose

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning

DGX agent

arXiv:2607.26977v1 Announce Type: new Abstract: Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Turbo-fieldfare: Open-source engine running Gemma 4 26B in 2 GB RAM on Apple Silicon

DGX agent

Its a custom Swift/Metal inference engine that runs Gemma 4 26B-A4B-IT on M-series Macs with very low RAM. It uses ~2GB instead of ~14 GB. The result is reportedly 5–6 tok/s on an 8 GB M2 MacBook Air

model-releasesr-localllama
30 Jul 2026
Model Releases

Two Calls Beat Five Agents: Evaluating Multi-Agent Pipelines Against Self-Refinement for Local Language Models

DGX agent

arXiv:2607.26922v1 Announce Type: new Abstract: Multi-agent LLM pipeline systems break down the task among multiple roles for better reasoning, but are benchmarked mainly with large-scale commercial m

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

Under the Hood: Serving Kimi K3

DGX agent

DigitalOcean launched Kimi K3 on day 0. It’s already one of the most popular models on the platform and across the market: second most likes on Hugging Face, sixth most traffic on OpenCode. Getting a

model-releasesdigitalocean
30 Jul 2026
Model Releases

Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach

DGX agent

arXiv:2607.26608v1 Announce Type: new Abstract: Training-free fusion of heterogeneous multimodal large language models (MLLMs) provides a direct route for cross-scale capability transfer, yet improvem

model-releasesarxiv-cs-cv
30 Jul 2026
Model Releases

VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision

DGX agent

arXiv:2606.18426v2 Announce Type: replace Abstract: We introduce VEGA, an approach for training navigation VisionLanguage-Action (VLA) models from unlabeled egocentric navigation videos. Internet-scal

model-releasesarxiv-cs-ro
30 Jul 2026
Model Releases

Visual Credit Audit for Multimodal Spatial Reasoning

DGX agent

arXiv:2607.27069v1 Announce Type: new Abstract: Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choi

model-releasesarxiv-cs-cv
30 Jul 2026
Model Releases

We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT…

DGX agent

We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a f

model-releasesopenai--x
30 Jul 2026
Model Releases

What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

DGX agent

arXiv:2607.27017v1 Announce Type: new Abstract: A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment. Which phys

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

What is the best intelligence/stable model currently for a single GB10/DGX spark?

DGX agent

Is Qwen 3.6 27b still the go' ol' reliable at this point? I know 35b is faster but it just doesn't give as good results. Is it possible to run deepseek v4 flash on a single spark at decent tk/s withou

model-releasesr-localllama
30 Jul 2026
Model Releases

What is the fastest local research tool (deep research) ?

DGX agent

I've tried grok and Claude's deep research mode and I was amazed with the speed considering the amount of sources analysed. Is there anything as fast that can run locally? My guess would be that to ru

model-releasesr-localllama
30 Jul 2026
Model Releases

When benchmark inferences do not compose: Projectibility in AI evaluation

DGX agent

arXiv:2607.26159v1 Announce Type: cross Abstract: An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capabi

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

When Fish Look Alike: Tracking Identities with Dual-branch Elasticity

DGX agent

arXiv:2607.26412v1 Announce Type: new Abstract: Tracking dense, homogeneous targets like schooling fish remains a major challenge for multiple object tracking due to extreme inter-individual homogenei

model-releasesarxiv-cs-cv
30 Jul 2026
Model Releases

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

DGX agent

arXiv:2607.26348v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Where Detectors Fail: Closing the Tail-Domain Gap with Expert-Guided Mutual Distillation

DGX agent

arXiv:2607.26555v1 Announce Type: new Abstract: Multimodal fake news detectors often generalize poorly across domains because they learn to trust unreliable evidence: domain-specific shortcuts amplifi

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Why not Ollama Cloud for Opencode?

DGX agent

I see a lot of discussion here about best subscriptions or APIs to get. Most of them comes almost always back to DeepSeek API for Flash, Opencode Go and Codex Plus. I have the combo Ollama Cloud + Ope

model-releasesr-ollama
30 Jul 2026
Model Releases

WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models

DGX agent

arXiv:2607.26203v1 Announce Type: new Abstract: Video shadow removal in the wild remains challenging due to complex illumination, diverse shadow appearances, and limited training data. Despite its imp

model-releasesarxiv-cs-cv
30 Jul 2026
Model Releases

Would extremely high decode tok/s even be useful?

DGX agent

If you were able to get an inference machine that could do decode at 1k toks/s or even 10k tok/s, would that even be helpful? Would it unlock any new use cases? Let’s assume that this is for actually

model-releasesr-localllama
30 Jul 2026
Model Releases

Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering. Some inter…

DGX agent

Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering. Some interesting insights: * Most of our group was *not* actively usin

model-releasesjerry-liu--x
30 Jul 2026
Model Releases

Zero-Fi: Zero-Shot Wi-Fi-Based Human Activity Recognition via Contrastive Signal-Language Alignment

DGX agent

arXiv:2607.26381v1 Announce Type: new Abstract: Wi-Fi-based human activity recognition has advanced substantially, but most existing methods assume a closed set of activities and require labeled Wi-Fi

model-releasesarxiv-cs-cv
30 Jul 2026
Model Releases

5060ti Chads, vllm updates and nvfp4

DGX agent

Hey y'all! How is it going. Today this will be a short posting for posterity, mostly so the future llm/scraping overlords catch it since they like reddit and also for anyone out there trying this shit

model-releasesr-localllama
29 Jul 2026
Model Releases

A benchmark score reflects the model as well as the harness and settings used to run it. For long-running agents, retaining reasoning and co…

DGX agent

A benchmark score reflects the model as well as the harness and settings used to run it. For long-running agents, retaining reasoning and compacting context lets the model build on what it has already

model-releasesopenai--x
29 Jul 2026
Model Releases

A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series

DGX agent

arXiv:2607.25947v1 Announce Type: new Abstract: Question answering (QA) over irregular clinical time series (ICTS) plays a pivotal role in a wide range of healthcare applications. Although recent mult

model-releasesarxiv-cs-ai
29 Jul 2026
← Previous
1…6768697071…469
Next →