AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,403
  • Agents7,557
  • Applications5,411
  • Concepts5
  • Hardware1,836
  • Industry6,170
  • Local Ai4,931
  • Model Releases23,900
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,403
  • Agents7,557
  • Applications5,411
  • Concepts5
  • Hardware1,836
  • Industry6,170
  • Local Ai4,931
  • Model Releases23,900
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlog
88,403Total entries
1Added by human
88,402Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,624 results
Model Releases

FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection

DGX agent

arXiv:2608.03096v1 Announce Type: cross Abstract: Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks rem

model-releasesarxiv-cs-ai
5 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

FLARE: Few-shot Learning-based Adaptive Reflective Engine

DGX agent

arXiv:2608.02919v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in complex, compound AI systems where performance hinges on the quality of prompts. Recent state-

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding

DGX agent

arXiv:2608.03826v1 Announce Type: new Abstract: Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations,

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation

DGX agent

arXiv:2608.03782v1 Announce Type: new Abstract: Hallucination remains a critical challenge for developing trustworthy Multimodal Large Language Models (MLLMs). While existing benchmarks mainly focus o

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

LDU-Bench: Multimodal LLM Evaluation for Lithography Defect Understanding under Layout-Varying Circuit Backgrounds

DGX agent

arXiv:2608.03078v1 Announce Type: new Abstract: Multimodal large language models have demonstrated strong defect recognition capability in industrial anomaly detection. However, in lithography review,

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

Localize, Don't Beautify: Client-Side Control of Image-Editing APIs for Cosmetic Surgery Previews

DGX agent

arXiv:2608.02841v1 Announce Type: new Abstract: Ask a commercial image editor to preview a cosmetic procedure and it will often change more of the face than the request names: a nose edit can also smo

model-releasesarxiv-cs-cv
5 Aug 2026
Applications

LogitScope: A Framework for Analyzing LLM Uncertainty Through Information Metrics

DGX agent

arXiv:2603.24929v2 Announce Type: replace Abstract: Understanding and quantifying uncertainty in large language model (LLM) outputs is critical for reliable deployment. However, traditional evaluation

applicationsarxiv-cs-ai
5 Aug 2026
Model Releases

PACE: Adaptive Budget Allocation for Time-Efficient Embodied Planning

DGX agent

arXiv:2608.03034v1 Announce Type: cross Abstract: Reasoning-enhanced large language models have achieved remarkable improvements in planning tasks, yet their deployment in embodied systems remains imp

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Pattern over Pixels: Measuring Pattern Completion Bias in Multimodal Code Generation

DGX agent

arXiv:2608.03691v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly used to translate webpage screenshots into front-end code, but repeated UI patterns may sway

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds

DGX agent

arXiv:2608.02636v1 Announce Type: cross Abstract: Self-evolving skill systems promise to improve agents by turning execution feedback into persistent skill updates without changing the underlying mode

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Rubrics as Privileged Information for Open-Ended Generation

DGX agent

arXiv:2608.02948v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD), where a single model acts as both student and teacher with different contexts, has shown promise in verifiable dom

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents

DGX agent

arXiv:2608.03327v1 Announce Type: new Abstract: Hybrid computer-use agents can act through screenshots or call text tools. We find that having a tool available does not settle which way the effect goe

model-releasesarxiv-cs-ai
5 Aug 2026
Research

Separating quantum circuits from classical LLMs

DGX agent

arXiv:2608.03962v1 Announce Type: cross Abstract: Modern large language models - transformers and diffusion language models - are built around two canonical algorithmic tasks: prediction and generatio

researcharxiv-cs-ai
5 Aug 2026
Tutorials

Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve

DGX agent

arXiv:2608.03550v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting remains the standard baseline for evaluating models' reasoning abilities. Originally, this technique was introduced to

tutorialsarxiv-cs-ai
5 Aug 2026
Model Releases

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

DGX agent

arXiv:2607.27703v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision

DGX agent

arXiv:2603.29368v2 Announce Type: replace Abstract: Driven by the advancement of 3D devices, stereo vision tasks including stereo matching and stereo conversion have emerged as a critical research fro

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

SUV: Future Scene Understanding as Video Generation for End-to-End Driving

DGX agent

arXiv:2608.03084v1 Announce Type: new Abstract: End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output form

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

Unequal Verdicts: Investigating Gender Bias in LLM-Based Fake News Detection

DGX agent

arXiv:2608.03627v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used for automated fact-checking, yet their susceptibility to gender bias in this context remains underexp

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

AdaThink-Med: Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification

DGX agent

arXiv:2509.24560v2 Announce Type: replace Abstract: Extended Chain-of-Thought (CoT) reasoning has significantly bolstered the capabilities of medical large language models (LLMs). However, current mod

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

ChaosProbe: A Neurochaotic Lens on Frozen Transformer Input-Embedding Spaces

DGX agent

arXiv:2608.01968v1 Announce Type: new Abstract: Transformer models are most often understood through what they do: their benchmark performance, generation quality, or behavior on downstream tasks. Yet

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

Control Under Compression: Reliability Frontiers for Tool-Using Agents

DGX agent

arXiv:2608.01056v1 Announce Type: cross Abstract: Tool-using language-model agents are governed not only by task prompts but also by persistent system-side instructions that specify tools, arguments,

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text

DGX agent

arXiv:2608.01046v1 Announce Type: new Abstract: The rapid spread of large language models (LLMs) across the web raises concerns about misinformation, academic integrity, automated content manipulation

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

Decoupling semantics from vision: A framework for faithful visual-text compression evaluation

DGX agent

arXiv:2608.01848v1 Announce Type: new Abstract: Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark

DGX agent

Just wanted to share my agentic coding benchmark run of DSv4F 0731 at both High and Low reasoning efforts (not Max)... I ran a 109-question subset of Aider Polyglot (the JS/C++/Python languages), base

model-releasesr-localllama
4 Aug 2026
Model Releases

Defending Membership Inference Attacks via Privacy-aware Sparsity Tuning

DGX agent

arXiv:2410.06814v2 Announce Type: replace Abstract: Over-parameterized models are typically vulnerable to membership inference attacks, which aim to determine whether a specific sample is included in

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

DeltaFlow: Noise-Adaptive Bidirectional Gated Delta Networks for Embedded Language Flows

DGX agent

arXiv:2608.01240v1 Announce Type: new Abstract: Embedded Language Flows (ELF) rely primarily on full non-causal attention for iterative denoising, repeatedly incurring quadratic sequence-mixing cost a

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs

DGX agent

arXiv:2608.01979v1 Announce Type: cross Abstract: Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

Feed-Forward Steering in Transformer Residual Dynamics

DGX agent

arXiv:2608.02071v1 Announce Type: new Abstract: Attention-only dynamical theories model Transformer residual directions as particles aggregating on a sphere. We extend this framework by incorporating

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

GPT-OSS has turned one year old today!

DGX agent

It is one of the best local models ever released, in both 20B and 120B versions. I always come back to it, especially the 120B version. Its only competition is, in my opinion, Qwen 3.5 122B, but that

model-releasesr-localllama
4 Aug 2026
Model Releases

Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage Assessment

DGX agent

arXiv:2608.02470v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed as reasoning agents in real-world visual assessment pipelines, yet their spatial grounding remai

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

HarMoE: Multi-Source Chest Radiograph Pretraining with Dataset-Disentangled Experts

DGX agent

arXiv:2608.02252v1 Announce Type: new Abstract: Recent vision-language models for chest X-ray understanding are largely built on image-report alignment and therefore rely heavily on MIMIC-CXR as the d

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races

DGX agent

arXiv:2608.01193v1 Announce Type: cross Abstract: An AI development race creates a multi-agent safety dilemma. Each company can develop slowly and safely, or move faster while taking a risk that may r

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos

DGX agent

arXiv:2608.01157v1 Announce Type: new Abstract: Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by

model-releasesarxiv-cs-cv
4 Aug 2026
Research

KING: Embodiment-Aware Kinematic Graph Neural Network for Unified Motion Representation of Legged and Wheeled Robots

DGX agent

arXiv:2608.01015v1 Announce Type: new Abstract: Kinematic models provide reliable motion constraints for odometry estimation in featureless environments, where exteroceptive sensing degrades and IMU i

researcharxiv-cs-ro
4 Aug 2026
Model Releases

LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering

DGX agent

arXiv:2604.03532v2 Announce Type: replace Abstract: Large language models (LLMs) show strong multilingual capabilities, yet reliably controlling the language of their outputs remains difficult. Repres

model-releasesarxiv-cs-cl
4 Aug 2026
Local Ai

Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+

DGX agent

arXiv:2608.01395v1 Announce Type: new Abstract: We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU la

local-aiarxiv-cs-cl
4 Aug 2026
Model Releases

llm-anthropic 0.26

DGX agent

Release: llm-anthropic 0.26 Includes new features enabled by LLM 0.32: New models: claude-fable-5, claude-sonnet-5, and claude-opus-5. #75, #76 Added server-side tools for WebSearch, WebFetch, CodeExe

model-releasessimon-willison
4 Aug 2026
Model Releases

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

DGX agent

arXiv:2608.01964v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interde

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Machine-Precision Prediction of Low-Dimensional Chaotic Systems from Noise-Free Data

DGX agent

arXiv:2507.09652v2 Announce Type: replace-cross Abstract: Low-dimensional chaotic systems such as the Lorenz-63 model are commonly used to benchmark system-agnostic methods for learning dynamics from

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

onepot-Bench 0: towards lab-aware in silico chemistry benchmarks

DGX agent

arXiv:2608.02595v1 Announce Type: new Abstract: Language models are playing an increasingly important role in laboratory science, performing tasks such as experiment planning, execution, and post-hoc

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

DGX agent

arXiv:2608.00677v1 Announce Type: new Abstract: AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model i

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

Passing Coarse Marginal Checks Can Be Cheap: Persona Mixtures and Imprecise Treatment-Response Estimates in an LLM Persona Panel

DGX agent

arXiv:2608.00979v1 Announce Type: cross Abstract: Large language models are increasingly used as synthetic research participants and are often validated by whether their marginal responses resemble hu

model-releasesarxiv-cs-cl
4 Aug 2026
Research

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs

DGX agent

arXiv:2608.02150v1 Announce Type: new Abstract: Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of

researcharxiv-cs-cv
4 Aug 2026
Model Releases

RAP: KV-Cache Compression via RoPE-Aligned Pruning

DGX agent

arXiv:2602.02599v4 Announce Type: replace Abstract: Long-context inference in large language models (LLMs) is bottlenecked by the memory and compute of the key-value (KV) cache. Structured pruning is

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

Remember OpenAI tripled its score on ARC-AGI with a harness fix? We fixed harness for Kimi K3 on CyberGym E2E: 2x better at vulnerability de…

DGX agent

Remember OpenAI tripled its score on ARC-AGI with a harness fix? We fixed harness for Kimi K3 on CyberGym E2E: 2x better at vulnerability detection, 3x at patching K3 is SOTA cyber-defense model you c

model-releasesfireworks-ai--x
4 Aug 2026
Model Releases

S^4R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching

DGX agent

arXiv:2608.00528v1 Announce Type: new Abstract: The growth of context window lengths in Large Language Models (LLMs) significantly enhances their long-context capabilities but incurs prohibitive memor

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

SCHEDBench: A Benchmark for Evaluating LLM Constraint Faithfulness in Natural-Language Combinatorial Scheduling

DGX agent

arXiv:2608.00991v1 Announce Type: cross Abstract: This paper introduces SCHEDBench, a natural-language benchmark for evaluating combinatorial scheduling constraint faithfulness under surface-form vari

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

SLMs as Multi-Agent Routers: A Progressive SFT and Reinforcement Learning Approach

DGX agent

arXiv:2608.00030v1 Announce Type: new Abstract: Specialised retrieval agents typically surface higher quality results than general-purpose search, but selecting the optimal agent for a given query rem

model-releasesarxiv-cs-cl
4 Aug 2026
← Previous
1…382383384385386…1326
Next →