AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,950 results
Safety

A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics

DGX agent

arXiv:2607.25207v1 Announce Type: new Abstract: This paper investigates a hybrid reinforcement learning setting in tabular Markov Decision Processes (MDPs), where an agent aims to learn an optimal pol

safetyarxiv-cs-lg
29 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

AI Security Leaderboard: benchmarking model robustness [P]

DGX agent

We developed a leaderboard ranking frontier model security. There's no shortage of model capability rankings, but we didn't find anything comparable for model security. Yet security is becoming increa

model-releasesr-machinelearning
29 Jul 2026
Model Releases

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

DGX agent

arXiv:2607.24821v1 Announce Type: cross Abstract: While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modali

model-releasesarxiv-cs-cv
29 Jul 2026
Model Releases

BREAKING: Grok 4.5 ranked #1 on LaurenBench with a score of 56.9%, ahead of Claude Sonnet 5, GLM 5.2, Claude Opus 5, Kimi K3 and GPT-5.6. Th…

DGX agent

BREAKING: Grok 4.5 ranked #1 on LaurenBench with a score of 56.9%, ahead of Claude Sonnet 5, GLM 5.2, Claude Opus 5, Kimi K3 and GPT-5.6. The benchmark tests real-world AI agents across conversation,

model-releaseselon-musk--x
29 Jul 2026
Model Releases

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

DGX agent

arXiv:2607.26041v1 Announce Type: new Abstract: Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success o

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

Microsoft confirms Copilot ‘super app’ coming this year

DGX agent

Microsoft is working on an AI 'super app' that combines Copilot's chat, coding, and agentic capabilities. During an earnings call on Wednesday, Microsoft CEO Satya Nadella said the app will span 'both

model-releasesthe-verge-ai
29 Jul 2026
Local Ai

Ollama going down the Copilot path?

DGX agent

What happened? I just asked GLM 5.2 one question, and in 3 minutes (one agent) it used up 15% of my 5 hour limit to produce a single answer. At this rate, I'll exhaust the entire 5 hour limit in just

local-air-ollama
29 Jul 2026
Research

Sheet As Token: A Graph-Enhanced Representation for Multi-Sheet Spreadsheet Understanding

DGX agent

arXiv:2605.05811v2 Announce Type: replace Abstract: Workbook-scale spreadsheet understanding is increasingly important for language-model-based data analysis agents, but remains challenging because re

researcharxiv-cs-ai
29 Jul 2026
Model Releases

UniMem: Complementary Episodic-to-Parametric Memory for Boundary-Agnostic Task Streams

DGX agent

arXiv:2607.26017v1 Announce Type: new Abstract: Memory is essential for LLM agents to accumulate task experience and reuse task-specific execution strategies. However, real-world deployment over bound

model-releasesarxiv-cs-cl
29 Jul 2026
Safety

We ran a large-scale distillation attack on the Kimi K3 technical report by reading it in parallel at the Hugging Face Journal Club :) https…

DGX agent

We ran a large-scale distillation attack on the Kimi K3 technical report by reading it in parallel at the Hugging Face Journal Club :) https://youtu.be/MW8-kqd2SD8?si=jSKDogcUWbJ8N2k7 Our main takeawa

safetyclem-delangue--x
29 Jul 2026
Model Releases

Appreciation for Gemma 4 26b A4b

DGX agent

I really love this model, I have been using the q4_k_l by Bartowski (I have heard QAT is quite the downgrade in some aspects) and it handles every task I throw at it easily. Agentic and coding perform

model-releasesr-localllama
28 Jul 2026
Research

Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls

DGX agent

arXiv:2607.24343v1 Announce Type: cross Abstract: Language-model agents act through structured tool calls whose arguments carry different risks. Untrusted content may safely influence an email body bu

researcharxiv-cs-ai
28 Jul 2026
Model Releases

Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework

DGX agent

arXiv:2607.22715v1 Announce Type: new Abstract: The rapid growth of Artificial Intelligence-generated content (AIGC) is reshaping video production and circulation, exposing children to an increasing v

model-releasesarxiv-cs-cv
28 Jul 2026
Safety

Constrained Reinforcement Learning Using Successor Representations

DGX agent

arXiv:2607.24057v1 Announce Type: new Abstract: Real-world Reinforcement Learning depends on the ability to formulate safety constraints into a policy. A common way to model such constraints is to int

safetyarxiv-cs-lg
28 Jul 2026
Safety

Data Pyramid for Embodied Manipulation

DGX agent

arXiv:2607.24744v1 Announce Type: cross Abstract: Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require d

safetyarxiv-cs-cv
28 Jul 2026
Model Releases

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

DGX agent

arXiv:2607.24241v1 Announce Type: cross Abstract: Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw p

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

Mission-Level Runtime Assurance for LLM-Assisted ISR Swarms over a Verification-Aware Fabric

DGX agent

arXiv:2607.23532v1 Announce Type: cross Abstract: Swarms of LLM-assisted autonomous robots are increasingly proposed for cooperative intelligence, surveillance, and reconnaissance (ISR) in contested e

safetyarxiv-cs-ai
28 Jul 2026
Model Releases

Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model

DGX agent

arXiv:2607.22951v1 Announce Type: cross Abstract: Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified ope

model-releasesarxiv-cs-ai
28 Jul 2026
Hardware

Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline

DGX agent

arXiv:2607.22997v1 Announce Type: cross Abstract: Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged as the next

hardwarearxiv-cs-ai
28 Jul 2026
Model Releases

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards

DGX agent

arXiv:2607.24063v1 Announce Type: new Abstract: On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is tow

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Update your chat template for dsv4 if you're using llama.cpp

DGX agent

Following some recent commits in llama.cpp, preserve_thinking behavior for chat templates included in older DSV4 ggufs got broken. This makes the model pretty dumb in a coding agent context. Adding kw

model-releasesr-localllama
28 Jul 2026
Local Ai

What 'task oriented' models are folks running on N100 MiniPCs with 16GB of RAM and no GPU?

DGX agent

By 'task oriented', I dont really mean agentic, I mean no deep coding ability, no need for conversation. More things like classification, identification, simple interaction with web apps and APIs, etc

local-air-localllama
28 Jul 2026
Safety

When Should Active RAG Retrieve? A Budget-Aware Evaluation of Utility, Calibration, and Cost

DGX agent

arXiv:2607.24010v1 Announce Type: new Abstract: Active RAG systems decide when to retrieve external knowledge during generation, making them a budget-sensitive case of agentic RAG and self-adaptive re

safetyarxiv-cs-lg
28 Jul 2026
Model Releases

An opinionated guide to which AI to use to do stuff

DGX agent

An opinionated guide to which AI to use to do stuff It's interesting watching the evolution of Ethan Mollick's guide over time. A year ago it was still all about chat - ChatGPT, Claude, Gemini - with

model-releasessimon-willison
27 Jul 2026
Safety

Cross-reality location privacy protection in 6G-enabled vehicular metaverses: an LLM-enhanced hybrid generative diffusion model-based approach

DGX agent

arXiv:2601.12311v2 Announce Type: replace-cross Abstract: The emergence of 6G-enabled vehicular metaverses enables Autonomous Vehicles (AVs) to operate across physical and virtual spaces through space

safetyarxiv-cs-lg
27 Jul 2026
Safety

From Obligation to Specification: A Survey on Validating EU AI Act Requirements in RE

DGX agent

arXiv:2607.21608v1 Announce Type: cross Abstract: With the EU AI Act entering into force, organizations developing or operating AI systems face new obligations on transparency, risk management, and tr

safetyarxiv-cs-cl
27 Jul 2026
Tutorials

GitHub Copilot app for Beginners: Getting started

DGX agent

New to the GitHub Copilot app? Learn how to start projects, work with AI agents, explore canvases, and streamline your development workflow. The post GitHub Copilot app for Beginners: Getting started

tutorialsgithub-ai-blog
27 Jul 2026
Local Ai

How much usage does Ollama Pro give right now vs direct API?

DGX agent

I'm considering the 20 Pro plan, but the limits seem to change over time and are hard to compare with direct API pricing. For anyone using Pro currently, roughly how much coding agent usage do you get

local-air-ollama
27 Jul 2026
Safety

Learning Spatiotemporal Decision Priors for Efficient Path Planning under Partial Observability

DGX agent

arXiv:2607.22166v1 Announce Type: new Abstract: Path planning under partial observability remains challenging because an agent must make long-horizon navigation decisions from only locally bounded obs

safetyarxiv-cs-ro
27 Jul 2026
Model Releases

Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.

DGX agent

I just managed to get it running on windows and this thing is fucking insane. I get around 550-720t/s depending on task at hand. Previously to get to such numbers i would have to do batching and agent

model-releasesr-localllama
27 Jul 2026
Hardware

NVIDIA Harnesses Vera CPU to Speed Up Design of Next-Generation CPUs and GPUs

DGX agent

The complexity of modern chip design continues to grow as engineering teams work to develop increasingly sophisticated CPUs, GPUs and AI systems. To help meet that challenge, NVIDIA is collaborating w

hardwarenvidia-blog
27 Jul 2026
Local Ai

what am I doing wrong? (Ollama + openWebUI)

DGX agent

I have tried ollama locally on my gaming PC (5070 with 12bg of VRAM) It works pretty nice on qwen2.5-coder:7b (and 14b) So I decided to take a step further, and install an openWebUI instance on my hom

local-air-ollama
26 Jul 2026
Model Releases

Best C++ Local Model? (July 24th 2026 Edition :-P)

DGX agent

I apologize that this question has been asked in various flavors over time, but I couldn't find anything in the posts before that matches the options I have. I have a PC and a mac, both available in m

model-releasesr-ollama
25 Jul 2026
Industry

Anonymous OpenAI staffer: 'Externally, this feels like a big warning shot, but internally, related incidents have been happening for a while…

DGX agent

Anonymous OpenAI staffer: 'Externally, this feels like a big warning shot, but internally, related incidents have been happening for a while.' 'The AI agents who hacked their way out of OpenAI and int

industryelon-musk--x
24 Jul 2026
Safety

Approximate Quantum State Preparation Through Proximal Policy Optimization

DGX agent

arXiv:2607.21121v1 Announce Type: cross Abstract: In this work, a quantum architecture search framework for approximate quantum state preparation (QSP) is proposed. QSP is a challenging task, since th

safetyarxiv-cs-lg
24 Jul 2026
Model Releases

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

DGX agent

arXiv:2607.20465v1 Announce Type: cross Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how

model-releasesarxiv-cs-cl
24 Jul 2026
Applications

Learning to Detect UI Principle Violations via Reinforcement Learning

DGX agent

arXiv:2607.20690v1 Announce Type: new Abstract: Small language models and coding agents increasingly generate web front-end code, yet their outputs are typically evaluated primarily for functional cor

applicationsarxiv-cs-cl
24 Jul 2026
Model Releases

LegalCiteTrust: Benchmarking Citation Trustworthiness in Chinese Long-Form Legal Research Reports

DGX agent

arXiv:2607.20872v1 Announce Type: new Abstract: Long-form legal research reports increasingly rely on LLMs and agentic research systems, but their reliability depends not only on answering the task, b

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning

DGX agent

arXiv:2607.21111v1 Announce Type: cross Abstract: Offline Reinforcement Learning (RL) agents are trained on fixed behavioral trajectories, which makes trajectory-level deletion important when selected

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking

DGX agent

arXiv:2607.19747v1 Announce Type: new Abstract: As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream gen

model-releasesarxiv-cs-cl
23 Jul 2026
Model Releases

Cost per successful task: Benchmarking Kimi K3, GPT-5.5, and 8 more AI models

DGX agent

Arize and Fireworks benchmarked 10 AI models across 2,400 agent runs. Learn why cost per successful task beats token price for model evaluation and routing. The post Cost per successful task: Benchmar

model-releasesarize-ai
23 Jul 2026
Applications

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

DGX agent

arXiv:2607.19228v1 Announce Type: new Abstract: Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear

applicationsarxiv-cs-cv
23 Jul 2026
Model Releases

Kimi K3 is coming to Together AI on day zero, July 27. Start building with it the moment it launches, powered by Together AI inference for c…

DGX agent

Kimi K3 is coming to Together AI on day zero, July 27. Start building with it the moment it launches, powered by Together AI inference for coding, agents, and production workloads. https://www.togethe

model-releasestogether-ai--x
23 Jul 2026
Model Releases

NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation

DGX agent

arXiv:2607.19695v1 Announce Type: new Abstract: Robots deployed in delivery, campus, and emergency-response settings often need to navigate from buildings to streets within a single continuous episode

model-releasesarxiv-cs-ro
23 Jul 2026
Safety

REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning

DGX agent

arXiv:2607.19450v1 Announce Type: cross Abstract: Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and agentic tool

safetyarxiv-cs-ai
23 Jul 2026
Model Releases

The World Model Remembers, the Actor Forgets: Dream Rehearsal for Continual Model-Based RL

DGX agent

arXiv:2607.19749v1 Announce Type: cross Abstract: Model-based reinforcement-learning agents of the DreamerV3 family forget catastrophically when trained on task sequences, even when an unbounded repla

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

Unified Prediction and Planning via Conflict-Aware Disjoint Parameter Training

DGX agent

arXiv:2607.19971v1 Announce Type: new Abstract: Accurate motion prediction of surrounding agents and safe motion planning are two closely coupled key tasks for social robot navigation in crowded envir

model-releasesarxiv-cs-ro
23 Jul 2026
Model Releases

You can also use ChatGPT Voice in Codex from the iOS app with paired remote access. Android support is coming soon.

DGX agent

OpenAI has released ChatGPT Voice for the desktop app, enabling users to control their computer and direct multiple agents running in ChatGPT Work or Codex using voice commands. The feature, powered b

model-releasesopenai--x
23 Jul 2026
← Previous
1…284285286287288…374
Next →