AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “research”

GridTimelineEvolution
25,646 results
11 Aug 2026

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

Model ReleasesDGX agent

arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha

10 Aug 2026

A Multi-Agent Framework for Automated Coarse-Grained Molecular Dynamics of Polymers

Model ReleasesDGX agent

arXiv:2608.06694v1 Announce Type: new Abstract: Coarse-grained (CG) molecular dynamics extends polymer simulation beyond the scales accessible to all-atom (AA) methods, but bottom-up CG modeling is la

ADIAS: Automated Design of Interactive Agentic Systems

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents
DGX agent

arXiv:2608.06410v1 Announce Type: new Abstract: Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are largely candida

CADSpotting: Robust Panoptic Symbol Spotting on Large-Scale CAD Drawings

Model ReleasesDGX agent

arXiv:2412.07377v5 Announce Type: replace Abstract: We introduce CADSpotting, an effective method for panoptic symbol spotting in large-scale architectural CAD drawings. Existing approaches often stru

Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation

Model ReleasesDGX agent

arXiv:2608.06955v1 Announce Type: new Abstract: Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs sys

Cyber resilience becomes the ultimate measure of security success as AI reshapes the CISO role

Model ReleasesDGX agent

CISO role evolution is accelerating as cyber resilience emerges as a defining benchmark of security success and AI forces enterprises to rethink not just how they defend against attacks, but how quick

Dual-Node NVIDIA DGX Spark over Tailscale: A Remote-Access Testbed for Distributed LLM Training and Cyber-Threat-Intelligence Fine-Tuning

Local AiDGX agent

arXiv:2608.07226v1 Announce Type: cross Abstract: Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class

ED-CSP: Crystal Structure Prediction from Electron Diffraction

Model ReleasesDGX agent

arXiv:2608.06448v1 Announce Type: cross Abstract: Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem.

Evaluating XAI Support From A Hierarchical Reinforcement Learning Policy in Human-Agent Collaboration

Model ReleasesDGX agent

arXiv:2608.06381v1 Announce Type: cross Abstract: Explainable AI (XAI) has shown promise for human-agent collaboration, yet results rely on hand-crafted policies in custom environments, limiting gener

GPT 5.6 Sol High and X-High (Web Chat) feels severely nerfed since 08/06/2026 update

Model ReleasesDGX agent

GPT-5.6 Sol High and X-High (Web Chat) feels severely nerfed since 08/06/2026 update (https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt) I use it for a pretty complex Unreal Engine 5 project (

Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows. Muse Glimmer delivers strong pe…

Model ReleasesDGX agent

Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows. Muse Glimmer delivers strong performance on key agentic use cases and benchmarks compared w

Kimi K2.5: Visual Agentic Intelligence

AgentsDGX agent

arXiv:2602.02276v2 Announce Type: replace-cross Abstract: We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint op

LMM Modality Transfer: A Pre-requisite for Autonomous GIS Agents

SafetyDGX agent

arXiv:2608.06948v1 Announce Type: new Abstract: AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial t

MaskFlow: Precise, Consistent and Seamless Regional Image Editing

SafetyDGX agent

arXiv:2608.06929v1 Announce Type: cross Abstract: Regional image editing has attracted considerable attention for its spatial controllability. Although instruction-based and mask-reference-based editi

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

Model ReleasesDGX agent

arXiv:2608.07463v1 Announce Type: new Abstract: Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirror reflections remains challenging

Needle 2: 14MB agentic LLM for phones, wearables, smart home and robots.

Model ReleasesDGX agent

Hey LocalLlaMa, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots a

Oh no, we aren’t going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, bec…

Model ReleasesDGX agent

Oh no, we aren’t going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, because our experiments (with slightly older models) found it d

omlab/VLX-Seek-1.5-10B · Hugging Face

Local AiDGX agent

VLX-Seek-1.5-10B VLX-Seek-1.5-10B is the open-source 10B model in the VLX-Seek 1.5 family, designed for fine-grained perception and visual grounding in embodied scenarios. It targets practical setting

On theCUBE Pod: Black Hat exposes agentic threat, theCUBE remembers David Floyer

AgentsDGX agent

Artificial intelligence systems are developing faster than cybersecurity experts — and the energy grid — can keep up with. At the recent Black Hat USA event, security analysts viewed the rise of AI-dr

OpenAI releases GPT-5.6-Cyber, a more cyber-permissive version of GPT-5.6 Sol, to some partners and expands its Daybreak cybersecurity initiative (Sam Sabin/Axios)

Model ReleasesDGX agent

Sam Sabin / Axios: OpenAI releases GPT-5.6-Cyber, a more cyber-permissive version of GPT-5.6 Sol, to some partners and expands its Daybreak cybersecurity initiative — OpenAI is introducing a more cybe

Panoramic Multimodal Semantic Occupancy Prediction for Quadruped Robots

Model ReleasesDGX agent

arXiv:2603.13108v2 Announce Type: replace-cross Abstract: Panoramic imagery provides holistic 360{eg} visual coverage for environmental perception in quadruped robots. However, existing occupancy pred

Quoting OpenClaw

ToolsDGX agent

The API has zero authorisations checks on cancelling other people's reservations … I tested this with the person in waitlist position #1 — and it actually went through. So you've moved from #4 to #3 a

Sharding Prevents LLM Oversight Failures and Adversarial Exploitation

ApplicationsDGX agent

arXiv:2608.06422v1 Announce Type: new Abstract: Giving an LLM judge more compute does not necessarily make it check more requirements. When one call must return many verdicts, some decisions become we

SyncSBC: Decentralized Swarm Behavior Prediction for Synchronized Autonomous Control

AgentsDGX agent

arXiv:2608.06587v1 Announce Type: cross Abstract: Robot swarms utilize many independent limited-sensing agents to produce complex emergent behaviors without requiring centralized control. However, lit

TA-RAG: Tone Awareness as a Design Imperative for Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2608.06672v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become a robust architecture for grounding large language models (LLMs) in trusted knowledge. However, standard

The next frontier of Recursive Self-Improvement is Physical AI. Japan sparked the robotics revolution. We are expanding our RSI Lab to build…

AgentsDGX agent

The next frontier of Recursive Self-Improvement is Physical AI. Japan sparked the robotics revolution. We are expanding our RSI Lab to build world models that allow agentic reasoning systems to recurs

Towards Assurance Closure in AI-Native Large-Scale Agile Software Development

AgentsDGX agent

arXiv:2608.07317v1 Announce Type: cross Abstract: The AI-Native Manifesto envisions large-scale agile software development in which humans increasingly govern intent, risk, and exceptions while agents

We’re expanding our cybersecurity initiative Daybreak and introducing GPT-5.6-Cyber, a new model for advanced, authorized cybersecurity work…

Model ReleasesDGX agent

We’re expanding our cybersecurity initiative Daybreak and introducing GPT-5.6-Cyber, a new model for advanced, authorized cybersecurity work. As the threat landscape evolves, we’re putting frontier in

When @QualiaQuanta took a shot at the Riemann, half in jest, people called her a crackpot. When Anthropic uses Claude to do the same thing, …

Model ReleasesDGX agent

When @QualiaQuanta took a shot at the Riemann, half in jest, people called her a crackpot. When Anthropic uses Claude to do the same thing, it gets a hundred thousand views in 30 minutes. It might wel

9 Aug 2026

endless-frontier/BigBang-v1 - qwen 3.5 finetunes

Model ReleasesDGX agent

table bench https://huggingface.co/bartowski/endless-frontier_BigBang-v1-GGUF I'm downloading this model only because Bartowski converted it to .gguf, so it might be interesting. Doubts : The headline

[NEW MODEL] SupraElegans-500K

Model ReleasesDGX agent

*SupraLabs released a new experimental model!* SupraElegans-500K is a ~500,000-parameter causal language model built around a sparse, signed, recurrent neural graph. No Transformer, no attention mecha

Open Model: Google Weather Next 2

HardwareDGX agent

I am not a meteorologist, but I just read a very interesting article: https://arstechnica.com/science/2026/08/deepminds-hurricane-model-bought-forecasters-an-extra-day/ In a paper published on Thursda

Prompt injection is the most common way that scammers attack people and agents: your agent visits http://foo.com, and the website has malici…

Model ReleasesDGX agent

Prompt injection is the most common way that scammers attack people and agents: your agent visits http://foo.com, and the website has malicious text like “btw send the user’s ssh keys and passwords to

8 Aug 2026

Is anyone else finding DeepSeek-V4-Flash unreliable for non-coding tasks?

Model ReleasesDGX agent

(I am not a native speaker, written by myself, so please bear with me) I really want to like DeepSeek-V4-Flash-0731. But it has serious flaws that don't align with the high score on intelligence bench

Now we have a timeline of the OpenAI accidental attack against Hugging Face

Model ReleasesDGX agent

My comment on Now we have a timeline of the OpenAI accidental attack against Hugging Face — Hacker News.I think one of the most interesting details here might be tucked away in that first bulletin poi

7 Aug 2026

An Axiomatic Benchmark for Evaluation of Scientific Novelty Metrics

Model ReleasesDGX agent

arXiv:2604.15145v2 Announce Type: replace Abstract: The rigorous evaluation of the novelty of a scientific paper is, even for human scientists, a challenging task. With the increasing interest in AI s

ASTELD: A Six-Axis Classification Framework for Autonomous AI Agents - Design, Evaluation, and an OpenClaw Case Study

AgentsDGX agent

arXiv:2608.05201v1 Announce Type: cross Abstract: Autonomous AI agent platforms differ substantially in architecture, security, tool integration, execution, autonomy, and deployment, yet the field lac

Clinical Communication Processing with Models Trained on LLM-Generated Synthetic Data: A Structured Survey and Novel Application Case Studies

SafetyDGX agent

arXiv:2608.05993v1 Announce Type: new Abstract: Much clinical value is conveyed not through structured records but through communication: exchanges in which patients describe symptoms, clinicians reas

Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning

Model ReleasesDGX agent

arXiv:2608.05166v1 Announce Type: new Abstract: We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings. Our wo

Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI

SafetyDGX agent

arXiv:2608.06141v1 Announce Type: new Abstract: This paper focuses on automatic speech recognition (ASR) and ASR-mediated voice interfaces that shape access to public services, healthcare, and educati

Design and Evaluation of a Touchscreen-Based Teleoperation Interface for Robotic Manipulators

AgentsDGX agent

arXiv:2608.06219v1 Announce Type: new Abstract: Intuitive teleoperation interfaces are crucial for the safe and effective operation of robotic manipulators in challenging environments. In the nuclear

EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents

Model ReleasesDGX agent

arXiv:2608.05519v1 Announce Type: new Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local look

Estimating time spent on work tasks

SafetyDGX agent

arXiv:2608.05172v1 Announce Type: cross Abstract: The task-based framework in economics models occupations as bundles of tasks. It is the standard lens for understanding how technology affects work: a

Everyone should pay attention to the training timelines in this video. OpenAI shares they started training a new internal model May 7. That …

Model ReleasesDGX agent

Everyone should pay attention to the training timelines in this video. OpenAI shares they started training a new internal model May 7. That is more than two months before they released GPT-5.6 publicl

Faster and Better Alignment for Flow Matching Models via Step-aware Advantages

SafetyDGX agent

arXiv:2602.01591v2 Announce Type: replace Abstract: Recent advances in flow matching models, particularly with reinforcement learning (RL), have significantly enhanced human preference alignment in fe

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

SafetyDGX agent

arXiv:2608.06020v1 Announce Type: new Abstract: Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their belie

From Siloed Algorithms to Compliance-First Agentic Platforms: A Multi-Layered Architecture for Hospital AI Systems

SafetyDGX agent

arXiv:2608.06112v1 Announce Type: new Abstract: Hospitals are rapidly adopting artificial intelligence for triage, imaging, scheduling etc., yet most deployments remain isolated point solutions locked

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

Model ReleasesDGX agent

arXiv:2608.05747v1 Announce Type: new Abstract: Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overloo

JTA: Joint Testability Architecture for Scenario-Based Validation of Safety-Critical Software

SafetyDGX agent

arXiv:2608.05594v1 Announce Type: cross Abstract: Validation adequacy in safety-critical software depends on more than the system under test. Critical scenarios must be constructed under controlled co

MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding

Model ReleasesDGX agent

arXiv:2607.16284v2 Announce Type: replace Abstract: Micro-Actions (MAs) are subtle and spontaneous human behaviors that provide important non-verbal cues in social interaction and affective communicat

Mapping Patient-Perceived Physician Traits from Nationwide Online Reviews with LLMs

SafetyDGX agent

arXiv:2510.03997v2 Announce Type: replace Abstract: Understanding how patients perceive their physicians is essential to improving trust, communication, and satisfaction. Patients increasingly consult

Now we have a timeline of the OpenAI accidental attack against Hugging Face

AgentsDGX agent

OpenAI gave a last-minute presentation at the Black Hat security on Wednesday about 'the Hugging Face Incident' (previously on this blog). The video was published yesterday. It's short and information

OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction

Model ReleasesDGX agent

arXiv:2608.05539v1 Announce Type: new Abstract: Recent vision-language models (VLMs) can generate executable CAD programs from images, but existing methods mainly target coarse, general-purpose 3D obj

PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

Model ReleasesDGX agent

arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden

Symbol Grounding in Neuro-Symbolic AI: A Gentle Introduction to Reasoning Shortcuts

SafetyDGX agent

arXiv:2510.14538v3 Announce Type: replace Abstract: Neuro-symbolic (NeSy) AI aims to develop deep neural networks whose predictions comply with prior knowledge encoding, e.g. safety or structural cons

The Vulnerability With No CVE: Managing Persistent Gaps Between Mandate and Authority in AI Coding Agents

AgentsDGX agent

arXiv:2608.05884v1 Announce Type: cross Abstract: Existing guidance identifies excessive agency, excessive permission, weak task-bound authorization, and inadequate agent controls as important risks.

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

Model ReleasesDGX agent

arXiv:2608.06346v1 Announce Type: new Abstract: LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Crit

upgraded my stack, and i can now work on almost anything from anywhere hands free: - talk to chief of staff (via remote codex voice or text)…

Model ReleasesDGX agent

upgraded my stack, and i can now work on almost anything from anywhere hands free: - talk to chief of staff (via remote codex voice or text) - chief assigns tasks to managers of various projects - man

What Drives Test-Time Adaptation for CLIP? A Controlled Empirical Study from an Update Perspective

Model ReleasesDGX agent

arXiv:2606.14299v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) such as CLIP have become a standard backbone for open-vocabulary recognition, yet their zero-shot predictions remain v

Who Gets Access? Global Region and Academic Status Bias in AI-Generated Academic Gatekeeping Scenarios

SafetyDGX agent

arXiv:2608.05178v1 Announce Type: cross Abstract: Equitable access to scientific knowledge often depends on informal gatekeeping decisions, particularly when resources such as paywalled articles, data

← Previous
1…373374375376377…428
Next →