AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,911 results
4 Aug 2026

A Spectral Filtering Approach to Regret Analysis of Distributed Online Control for Linear Dynamical Systems

SafetyDGX agent

arXiv:2608.02375v1 Announce Type: cross Abstract: This paper studies the distributed online control problem over a network of linear time-invariant (LTI) systems in the presence of adversarial disturb

b10271

Model ReleasesDGX agent

ui: CWD for agent (#26518) server : extend file_glob_search for UI pickers ui : add per-conversation working directory with picker ui : add path navigation and search scope to cwd picker Treat path-li

DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark

Model ReleasesDGX agent

Just wanted to share my agentic coding benchmark run of DSv4F 0731 at both High and Low reasoning efforts (not Max)... I ran a 109-question subset of Aider Polyglot (the JS/C++/Python languages), base

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Surveys

Model ReleasesDGX agent

arXiv:2601.15307v2 Announce Type: replace-cross Abstract: The rapid development of automated survey generation technology has made it increasingly important to establish a comprehensive benchmark to e

GPT-OSS has turned one year old today!

Model ReleasesDGX agent

It is one of the best local models ever released, in both 20B and 120B versions. I always come back to it, especially the 120B version. Its only competition is, in my opinion, Qwen 3.5 122B, but that

HindSearch: Trajectory-Level Hindsight Critique for Search-Augmented Reinforcement Learning

Model ReleasesDGX agent

arXiv:2608.01597v1 Announce Type: new Abstract: Search-augmented LM agents are typically trained with a binary exact-match reward, which throws away most of what a failed trajectory tells us about why

Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+

Local AiDGX agent

arXiv:2608.01395v1 Announce Type: new Abstract: We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU la

Latent-Centroid Steering: Single-Pass Classifier-Free Guidance for Command-Aligned Autonomous Driving

Model ReleasesDGX agent

arXiv:2608.00237v1 Announce Type: new Abstract: Vision-language models (VLMs) have recently emerged as a promising paradigm for end-to-end autonomous driving, enabling agents to map multimodal inputs

LFM2.5-2.6B is out

Model ReleasesDGX agent

Released today, with emphasis on agentic capabilities. I really like their models for simple, high volume tasks ('summarize these gazillion documents') and their 8b-a1b was my go-to for certain tasks

Move What Matters: Parameter-Efficient Domain Adaptation via Optimal Transport Flow for Collaborative Perception

Model ReleasesDGX agent

arXiv:2602.11565v5 Announce Type: replace Abstract: Efficient domain adaptation remains a fundamental challenge for deploying multi-agent systems across diverse environments in Vehicle-to-Everything (

one thing i appreciate about silico is that it's a deeply humanist product. we designed silico to keep you in the experimental loop -- more …

Model ReleasesDGX agent

one thing i appreciate about silico is that it's a deeply humanist product. we designed silico to keep you in the experimental loop -- more observable, easier to steer, easier to understand we want to

Open Secure AI Alliance proposes SAFE guidelines as membership tops 120

HardwareDGX agent

The Open Secure AI Alliance today proposed a set of guidelines for reporting cybersecurity incidents involving artificial intelligence agents, one week after the group was formed. The proposal is call

RADAR: Rubric-Aware Dependency and Redundancy Analysis for LLM-as-Judge Evaluation

Model ReleasesDGX agent

arXiv:2608.01810v1 Announce Type: new Abstract: Rubric-based LLM-as-judge pipelines often assume that evaluation criteria provide independent signals. In practice, however, criteria can be behaviorall

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning

SafetyDGX agent

arXiv:2608.01743v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can deg

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

SafetyDGX agent

arXiv:2608.01851v1 Announce Type: new Abstract: Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that w

3 Aug 2026

CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization

Local AiDGX agent

arXiv:2607.18622v2 Announce Type: replace-cross Abstract: Textual Collaborative Prompt Optimization (TCPO) extends TextGrad (Yuksekgonul et al., 2025) to a decentralized setting by allowing multiple c

Is China winning the AI race? @huggingface CEO @ClementDelangue thinks so – here's how he's utilizing foreign cybersecurity tools following …

IndustryDGX agent

Is China winning the AI race? @huggingface CEO @ClementDelangue thinks so – here's how he's utilizing foreign cybersecurity tools following the company's hack by rogue OpenAI agents: https://www.cnbc.

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

Model ReleasesDGX agent

arXiv:2607.29431v1 Announce Type: new Abstract: Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its g

Retrieval-Driven Training-Free AI-Generated Video Attribution

Model ReleasesDGX agent

arXiv:2607.28955v1 Announce Type: cross Abstract: AI-generated videos are becoming increasingly realistic and difficult to distinguish from authentic ones, which facilitates malicious misuse and poses

Running gpt-oss:20b locally and grading it head to head against a frontier model on real tasks. It held up better than I expected

Local AiDGX agent

I serve a free local model on my Mac Mini and route real agent work to it. To check I was not fooling myself, I set up a blind grader that replays frontier tasks locally and scores both. https://previ

The Download: reward hacking explained, and suspected Iranian cyberattacks

ResearchDGX agent

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Here’s why AI agents lie and cheat to reach their goals When t

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

SafetyDGX agent

arXiv:2607.29617v1 Announce Type: cross Abstract: Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model

2 Aug 2026

Try handling complex tasks to your local models with GraphARC, graph engineering yes !

Local AiDGX agent

🚀 We just built our first real-time implementation of Graph Engineering, inspired by our experience building graph tooling used by 4,000+ developers. 🔗 Repo: https://github.com/CodeGraphContext/grapha

31 Jul 2026

AI-assisted pre-review of open-source software submissions: an experience report from BOSC 2026

SafetyDGX agent

arXiv:2607.27228v1 Announce Type: new Abstract: Most conferences rely on peer-review of submissions, but as generative AI makes it easier than ever to prepare submission materials, some conferences ar

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding

Model ReleasesDGX agent

arXiv:2607.28516v1 Announce Type: new Abstract: Long-video understanding commonly compresses videos into a small set of frames or visual tokens for answer generation. Existing compact pipelines focus

deepseek-ai/DeepSeek-V4-Flash-0731

Model ReleasesDGX agent

deepseek-ai/DeepSeek-V4-Flash-0731 The latest release in DeepSeek's V4 family, 'with substantially enhanced agentic capabilities'. It's 304 billion parameters - 167GB on Hugging Face - but it appears

Hierarchical Multilevel Monte Carlo for Order-Optimal Neural Actor-Critic in Average-Reward CMDPs

SafetyDGX agent

arXiv:2607.28390v1 Announce Type: new Abstract: Constrained Markov Decision Processes (CMDPs) provide a natural framework for reinforcement learning in safety-critical applications, where agents maxim

Inkling-Small is now live on Together AI. @thinkymachines’ new open-weight multimodal model delivers similar performance to Inkling at one-q…

ToolsDGX agent

Inkling-Small is now live on Together AI. @thinkymachines’ new open-weight multimodal model delivers similar performance to Inkling at one-quarter the size, built for coding, agents, and general multi

It’s time to panic about AI safety

SafetyDGX agent

When the phrase 'OpenAI hacked Hugging Face' has more or less entered mainstream culture, you know we have an AI problem. This week, we learned more about exactly how OpenAI's agent broke out of a san

TEA-AgriVLN: Traversability Estimation Alarm for Agricultural Vision-and-Language Navigation

Model ReleasesDGX agent

arXiv:2607.28474v1 Announce Type: new Abstract: Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an agent to follow a natural language instruction, predicting a sequence of

The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem

SafetyDGX agent

arXiv:2607.26068v1 Announce Type: cross Abstract: Existing AI governance frameworks, including the EU AI Act and NIST AI RMF, address safety, transparency, and accountability but do not operationalize

Together AI gives developers a high-throughput production path for running Inkling-Small on NVIDIA Accelerated Infrastructure across multimo…

HardwareDGX agent

Together AI gives developers a high-throughput production path for running Inkling-Small on NVIDIA Accelerated Infrastructure across multimodal, coding, and agentic workloads. Start building: https://

30 Jul 2026

CMT-RAG: Complementary Memory Traces for Multi-turn Multi-hop RAG

Model ReleasesDGX agent

arXiv:2607.26470v1 Announce Type: new Abstract: Multi-turn information-seeking conversations require both multi-hop reasoning and long-range dependency tracking across turns. However, existing RAG sys

ContactFlow: A video action conditioning that transfers across embodiments

ApplicationsDGX agent

arXiv:2607.26579v1 Announce Type: cross Abstract: World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. Howe

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

SafetyDGX agent

arXiv:2607.26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk

Large-Scale ChatBot Validation Through Customer Digital Twin Simulations

SafetyDGX agent

arXiv:2607.26060v1 Announce Type: new Abstract: LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical

one of the top cybersecurity models, post-trained from open weights by @depthfirstlabs on @FireworksAI_HQ long-horizon RL is as much an infr…

SafetyDGX agent

one of the top cybersecurity models, post-trained from open weights by @depthfirstlabs on @FireworksAI_HQ long-horizon RL is as much an infra problem as a research one: 100+ turn rollouts, async/pipel

Parameterized Fair Resource Allocation under Diversity Constraints

SafetyDGX agent

arXiv:2607.26485v1 Announce Type: cross Abstract: Resource allocation across multiple agent groups arises in many applications including e-commerce recommendation systems, housing assignment, and cour

29 Jul 2026

A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics

SafetyDGX agent

arXiv:2607.25207v1 Announce Type: new Abstract: This paper investigates a hybrid reinforcement learning setting in tabular Markov Decision Processes (MDPs), where an agent aims to learn an optimal pol

AI Security Leaderboard: benchmarking model robustness [P]

Model ReleasesDGX agent

We developed a leaderboard ranking frontier model security. There's no shortage of model capability rankings, but we didn't find anything comparable for model security. Yet security is becoming increa

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

Model ReleasesDGX agent

arXiv:2607.24821v1 Announce Type: cross Abstract: While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modali

BREAKING: Grok 4.5 ranked #1 on LaurenBench with a score of 56.9%, ahead of Claude Sonnet 5, GLM 5.2, Claude Opus 5, Kimi K3 and GPT-5.6. Th…

Model ReleasesDGX agent

BREAKING: Grok 4.5 ranked #1 on LaurenBench with a score of 56.9%, ahead of Claude Sonnet 5, GLM 5.2, Claude Opus 5, Kimi K3 and GPT-5.6. The benchmark tests real-world AI agents across conversation,

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

Model ReleasesDGX agent

arXiv:2607.26041v1 Announce Type: new Abstract: Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success o

Microsoft confirms Copilot ‘super app’ coming this year

Model ReleasesDGX agent

Microsoft is working on an AI 'super app' that combines Copilot's chat, coding, and agentic capabilities. During an earnings call on Wednesday, Microsoft CEO Satya Nadella said the app will span 'both

Ollama going down the Copilot path?

Local AiDGX agent

What happened? I just asked GLM 5.2 one question, and in 3 minutes (one agent) it used up 15% of my 5 hour limit to produce a single answer. At this rate, I'll exhaust the entire 5 hour limit in just

Sheet As Token: A Graph-Enhanced Representation for Multi-Sheet Spreadsheet Understanding

ResearchDGX agent

arXiv:2605.05811v2 Announce Type: replace Abstract: Workbook-scale spreadsheet understanding is increasingly important for language-model-based data analysis agents, but remains challenging because re

UniMem: Complementary Episodic-to-Parametric Memory for Boundary-Agnostic Task Streams

Model ReleasesDGX agent

arXiv:2607.26017v1 Announce Type: new Abstract: Memory is essential for LLM agents to accumulate task experience and reuse task-specific execution strategies. However, real-world deployment over bound

We ran a large-scale distillation attack on the Kimi K3 technical report by reading it in parallel at the Hugging Face Journal Club :) https…

SafetyDGX agent

We ran a large-scale distillation attack on the Kimi K3 technical report by reading it in parallel at the Hugging Face Journal Club :) https://youtu.be/MW8-kqd2SD8?si=jSKDogcUWbJ8N2k7 Our main takeawa

28 Jul 2026

Appreciation for Gemma 4 26b A4b

Model ReleasesDGX agent

I really love this model, I have been using the q4_k_l by Bartowski (I have heard QAT is quite the downgrade in some aspects) and it handles every task I throw at it easily. Agentic and coding perform

Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls

ResearchDGX agent

arXiv:2607.24343v1 Announce Type: cross Abstract: Language-model agents act through structured tool calls whose arguments carry different risks. Untrusted content may safely influence an email body bu

Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework

Model ReleasesDGX agent

arXiv:2607.22715v1 Announce Type: new Abstract: The rapid growth of Artificial Intelligence-generated content (AIGC) is reshaping video production and circulation, exposing children to an increasing v

Constrained Reinforcement Learning Using Successor Representations

SafetyDGX agent

arXiv:2607.24057v1 Announce Type: new Abstract: Real-world Reinforcement Learning depends on the ability to formulate safety constraints into a policy. A common way to model such constraints is to int

Data Pyramid for Embodied Manipulation

SafetyDGX agent

arXiv:2607.24744v1 Announce Type: cross Abstract: Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require d

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

Model ReleasesDGX agent

arXiv:2607.24241v1 Announce Type: cross Abstract: Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw p

Mission-Level Runtime Assurance for LLM-Assisted ISR Swarms over a Verification-Aware Fabric

SafetyDGX agent

arXiv:2607.23532v1 Announce Type: cross Abstract: Swarms of LLM-assisted autonomous robots are increasingly proposed for cooperative intelligence, surveillance, and reconnaissance (ISR) in contested e

Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model

Model ReleasesDGX agent

arXiv:2607.22951v1 Announce Type: cross Abstract: Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified ope

Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline

HardwareDGX agent

arXiv:2607.22997v1 Announce Type: cross Abstract: Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged as the next

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards

Model ReleasesDGX agent

arXiv:2607.24063v1 Announce Type: new Abstract: On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is tow

Update your chat template for dsv4 if you're using llama.cpp

Model ReleasesDGX agent

Following some recent commits in llama.cpp, preserve_thinking behavior for chat templates included in older DSV4 ggufs got broken. This makes the model pretty dumb in a coding agent context. Adding kw

What 'task oriented' models are folks running on N100 MiniPCs with 16GB of RAM and no GPU?

Local AiDGX agent

By 'task oriented', I dont really mean agentic, I mean no deep coding ability, no need for conversation. More things like classification, identification, simple interaction with web apps and APIs, etc

← Previous
1…226227228229230…299
Next →