AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,737 results
Model Releases

CohereLabs/North-Micro-Vision-Instruct · Hugging Face

DGX agent

North Micro Vision Instruct is a 2.4B-parameter open-weight vision-language model with native-resolution image support, released under the Apache 2.0 license. It is designed as a compact foundation fo

model-releasesr-localllama
12 Aug 2026
Research

Do LLMs Benefit From Their Own Words?

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2602.24287v2 Announce Type: replace-cross Abstract: In multi-turn conversations, large language models typically condition on the full conversation history: both past user prompts and assistant

researcharxiv-cs-ai
12 Aug 2026
Model Releases

DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving

DGX agent

arXiv:2608.10413v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

E^3mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment

DGX agent

arXiv:2608.10796v1 Announce Type: new Abstract: Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interact

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation

DGX agent

arXiv:2608.10805v1 Announce Type: cross Abstract: Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's receptive fie

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models

DGX agent

arXiv:2608.10444v1 Announce Type: cross Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This pro

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes

DGX agent

arXiv:2608.10886v1 Announce Type: new Abstract: Robots operating in human environments need memories that capture not only what objects exist and where, but also how people use them over time and how

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

Google unveils the $399 Pixel Watch 5 with a satin pyrite case finish, offline Gemini, proactive AI suggestions, better GPS maps, and insulin resistance trends (Victoria Song/The Verge)

DGX agent

Victoria Song / The Verge: Google unveils the 399 Pixel Watch 5 with a satin pyrite case finish, offline Gemini, proactive AI suggestions, better GPS maps, and insulin resistance trends — The 399 Goog

model-releasestechmeme
12 Aug 2026
Safety

Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models

DGX agent

arXiv:2608.10393v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their adversarial r

safetyarxiv-cs-ai
12 Aug 2026
Model Releases

HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models

DGX agent

arXiv:2506.03922v4 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchma

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

HUI360: A 360{eg} Egocentric Dataset and Baselines for Human-Robot Interaction Anticipation

DGX agent

arXiv:2608.11051v1 Announce Type: new Abstract: As robots increasingly operate in human-populated environments, anticipating human intentions is essential for enabling proactive and socially aware beh

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

Idea for a deepseek-v4-flash-0731 backed automated research workflow to be leveraged via qwen3.6/3.8 27b for difficult tasks that require highly technical, not easy to find information.

DGX agent

Sometimes you have tasks that are outside of your expertise and the idea is this workflow automation could be leveraged to manage to have local AI figure it out using research from his workflow gather

model-releasesr-localllama
12 Aug 2026
Model Releases

imo GDPVal is probably the most important benchmark, it measures the performance of models on real world tasks Big leap in performance here …

DGX agent

imo GDPVal is probably the most important benchmark, it measures the performance of models on real world tasks Big leap in performance here to top it at a great price, congrats to @SpaceXAI team & loo

model-releasesemad-mostaque--x
12 Aug 2026
Safety

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

DGX agent

arXiv:2608.10635v1 Announce Type: cross Abstract: Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain chall

safetyarxiv-cs-ai
12 Aug 2026
Safety

Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting Logs

DGX agent

arXiv:2608.10050v1 Announce Type: new Abstract: Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business ch

safetyarxiv-cs-lg
12 Aug 2026
Model Releases

Optimal Stopping of Self-Refining Foundation Models

DGX agent

arXiv:2608.10729v1 Announce Type: cross Abstract: Foundation models can improve their outputs through a self-refinement process driven by external feedback. In this process, the model is embedded in a

model-releasesarxiv-cs-ai
12 Aug 2026
Research

Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences

DGX agent

arXiv:2608.10908v1 Announce Type: cross Abstract: As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the jud

researcharxiv-cs-cl
12 Aug 2026
Safety

Predicting Space Groups of Double Perovskites by LLM with Dynamic Few-Shot Learning

DGX agent

arXiv:2608.10483v1 Announce Type: new Abstract: Double perovskites (DPs) offer broad compositional tunability, but predicting the space groups (SGs) of stable structures remains difficult because avai

safetyarxiv-cs-ai
12 Aug 2026
Safety

Procedural Fairness Failures in RLHF from Preference Averaging

DGX agent

arXiv:2608.10126v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity. Wh

safetyarxiv-cs-ai
12 Aug 2026
Local Ai

Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent Memory

DGX agent

arXiv:2608.11066v1 Announce Type: cross Abstract: We prove inference-time quantum coordination advantages for specified AI state-tracking tasks. A solver compresses semantic history into a future-acce

local-aiarxiv-cs-ai
12 Aug 2026
Model Releases

Qwen3.8-Max is live on Together AI on day zero. Proud to have supported early testing alongside @Alibaba_Qwen, @vllm_project, and Inferact t…

DGX agent

Qwen3.8-Max is live on Together AI on day zero. Proud to have supported early testing alongside @Alibaba_Qwen, @vllm_project, and Inferact to get day-zero support live. We’re continuing to optimize pe

model-releasestogether-ai--x
12 Aug 2026
Local Ai

R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video

DGX agent

arXiv:2608.11017v1 Announce Type: cross Abstract: Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved, when it la

local-aiarxiv-cs-ai
12 Aug 2026
Model Releases

ReLTEx: Reliable LLM-based Taxonomy Expansion

DGX agent

arXiv:2608.10970v1 Announce Type: cross Abstract: Recent advances in Large Language Models (LLMs) have demonstrated strong capabilities in generating semantically relevant concepts and relations, maki

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

Rethinking Text-Based Image Retrieval in Specific Domain

DGX agent

arXiv:2608.10524v1 Announce Type: cross Abstract: Driven by the rapid advancement of vision-language representation learning, Text-based Image Retrieval (TBIR) has made notable progress. However, exis

model-releasesarxiv-cs-ai
12 Aug 2026
Safety

Scheduling Mixed RL Rollouts Beyond Prefix Locality

DGX agent

arXiv:2608.11152v1 Announce Type: cross Abstract: Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasingly combine rollout workloads across multiple dom

safetyarxiv-cs-lg
12 Aug 2026
Model Releases

Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72

DGX agent

Alibaba released the open‑weights Qwen3.8‑2.4T‑A95B (Qwen3.8‑Max), a fine‑grained mixture‑of‑experts model with 2.4 trillion parameters, hybrid full‑ and linear‑attention, a one‑million‑token context

model-releasesnvidia-developer
12 Aug 2026
Model Releases

SPACEXAI: Grok 4.6 leads on the two strongest knowledge-work / real-world productivity benchmarks (GDPVal-AA and AA-Briefcase) and on the le…

DGX agent

SPACEXAI: Grok 4.6 leads on the two strongest knowledge-work / real-world productivity benchmarks (GDPVal-AA and AA-Briefcase) and on the legal benchmark, while remaining highly competitive on coding-

model-releaseselon-musk--x
12 Aug 2026
Research

StreamFlow: Dynamic Memory Flows for Streaming Video Understanding

DGX agent

arXiv:2608.10949v1 Announce Type: cross Abstract: Streaming video understanding requires multimodal large language models (MLLMs) to preserve relevant evidence from continuously evolving streams under

researcharxiv-cs-cl
12 Aug 2026
Model Releases

The most dangerous document extraction failure isn't a wrong value. It's a missing row that looks like nothing is wrong. We released Extract…

DGX agent

The most dangerous document extraction failure isn't a wrong value. It's a missing row that looks like nothing is wrong. We released ExtractBench yesterday: 370 enterprise docs, 14 systems. The hardes

model-releasesjerry-liu--x
12 Aug 2026
Safety

Toward a Theory of Value in AI Alignment

DGX agent

arXiv:2608.10327v1 Announce Type: new Abstract: Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms

safetyarxiv-cs-ai
12 Aug 2026
Industry

Twitch streamers can now opt out from training Amazon’s AI

DGX agent

Twitch users can now opt out of allowing their content to be used to train Amazon's generative AI models. Opting out means that 'your streams, VODs, clips, stream chats, and pictures and text on your

industrythe-verge-ai
12 Aug 2026
Model Releases

When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning

DGX agent

arXiv:2608.09942v1 Announce Type: cross Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data

DGX agent

arXiv:2608.11022v1 Announce Type: cross Abstract: Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturing their con

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

1 Day in and I feel okay saying Muse-Glimmer-30B finally beats 3.6-27B for the size in some use-cases

DGX agent

A few things right off the bat: it reasons very efficiently. Like Grok 4.5 levels of efficient thinking it quantizes very well. My first few tests with iq3_xxs were better than Qwen/Gemma behaved at t

model-releasesr-localllama
11 Aug 2026
Research

A foundation model of numerical intelligence with cross-disciplinary generalization

DGX agent

arXiv:2607.28432v2 Announce Type: replace Abstract: Intelligence is commonly understood as the ability to acquire and apply knowledge, adapt to unfamiliar situations and solve new problems. Large lang

researcharxiv-cs-ai
11 Aug 2026
Model Releases

Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks

DGX agent

arXiv:2603.00077v3 Announce Type: replace-cross Abstract: Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduc

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making

DGX agent

arXiv:2608.07584v1 Announce Type: new Abstract: Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images. Many such t

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

Cross-Model Humor Preference Modeling with Cards Against Humanity

DGX agent

arXiv:2608.07481v1 Announce Type: cross Abstract: This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style

model-releasesarxiv-cs-ai
11 Aug 2026
Safety

DH-VLM: Dual-Horizon Cooperative Latent Reasoning for Autonomous Driving

DGX agent

arXiv:2608.09333v1 Announce Type: new Abstract: Large-scale language models for autonomous driving enable enhanced global understanding and long-horizon planning. However, when deployed in isolated ve

safetyarxiv-cs-ro
11 Aug 2026
Safety

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

DGX agent

arXiv:2608.09542v1 Announce Type: cross Abstract: Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent

safetyarxiv-cs-ai
11 Aug 2026
Safety

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition

DGX agent

arXiv:2608.09762v1 Announce Type: new Abstract: Real-world online reinforcement learning (RL) provides a promising approach for training robotic manipulation policies directly in the physical world, a

safetyarxiv-cs-ro
11 Aug 2026
Model Releases

ELICITED: EHR-grounded Longitudinal Interactive Conversations for Information-seeking Triage Evaluation and Decision-making

DGX agent

arXiv:2608.09024v1 Announce Type: new Abstract: Emergency-department (ED) triage requires clinicians to rapidly identify patients who need immediate attention, determine who can safely wait, and prior

model-releasesarxiv-cs-cl
11 Aug 2026
Research

Emotion in an active inference model of human driving

DGX agent

arXiv:2608.07480v1 Announce Type: new Abstract: Active inference has emerged as a principled framework for modeling adaptive behavior by balancing goal-directed action with uncertainty reduction. It h

researcharxiv-cs-ai
11 Aug 2026
Research

EvalConvoLearn: An Open-Source Framework for Evaluating Grounded Learner Simulations in Tutoring Conversations

DGX agent

arXiv:2608.07497v1 Announce Type: cross Abstract: Conversational learner simulations are valuable tools for testing learning theories, evaluating instructional materials and automated tutors, or power

researcharxiv-cs-cl
11 Aug 2026
Model Releases

ExtractBench is one of the most comprehensive benchmarks for real-world document extraction. ✅ It covers 4869 pages, across 67 document type…

DGX agent

ExtractBench is one of the most comprehensive benchmarks for real-world document extraction. ✅ It covers 4869 pages, across 67 document types, spanning 8 real-world domains: finance, energy, gov, auto

model-releasesjerry-liu--x
11 Aug 2026
Model Releases

FaLCon: Facet-Anchored Retrieval with Late Consensus for Sim2Real Text-Based Person Anomaly Search

DGX agent

arXiv:2608.09474v1 Announce Type: new Abstract: Text-based person anomaly search requires retrieving real-world pedestrian images from detailed natural-language descriptions using models trained prima

model-releasesarxiv-cs-cv
11 Aug 2026
Safety

Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective

DGX agent

arXiv:2608.08445v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is widely regarded as a novel paradigm born from the limitations of large language models (LLMs)--a mechanism to gr

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

From Diagnosis to Correction: Benchmarking and Improving Real-World Table Parsing

DGX agent

arXiv:2608.09842v1 Announce Type: new Abstract: Recent document parsers achieve table TEDS scores above 93 on OmniDocBench v1.6, yet community feedback and our audit reveal persistent failures on comp

model-releasesarxiv-cs-cv
11 Aug 2026
← Previous
1…325326327328329…370
Next →