AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
8 Jul 2026

SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models

AgentsDGX agent

arXiv:2512.18542v3 Announce Type: replace-cross Abstract: AI coding assistants produce vulnerable code in 45% of security-relevant scenarios~ite{veracode2025}, yet no public training dataset teaches b

VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery

AgentsDGX agent

arXiv:2607.06374v1 Announce Type: new Abstract: Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact e

You don't need heavyweight VLMs to OCR simple text-only PDFs. Doing that is like bringing a bazooka to a knife-fight, and is completely unne…

AgentsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

You don't need heavyweight VLMs to OCR simple text-only PDFs. Doing that is like bringing a bazooka to a knife-fight, and is completely unnecessary and worse quality than a tuned OCR approach. Output

7 Jul 2026

A Mathematical Theory of Value: a synthesis on goal-directed agency under resource constraints

SafetyDGX agent

arXiv:2606.12502v2 Announce Type: replace-cross Abstract: We propose that value -- the quantity goal-directed agents create, destroy, and exchange -- is a lawful structural quantity in the same catego

Conflict-Based Lazy Search for Fast Multi-Manipulator Planning

AgentsDGX agent

arXiv:2607.04124v1 Announce Type: cross Abstract: Employing multiple manipulators can boost efficiency and accomplish tasks that a single manipulator cannot do. However, real-time planning for multipl

Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI

Model ReleasesDGX agent

arXiv:2607.02873v1 Announce Type: cross Abstract: Large language model agents driving security tool suites over the Model Context Protocol are increasingly common. Yet the factors that bound their cap

Flow-A11y: Flow-Aware Accessibility Testing

AgentsDGX agent

arXiv:2607.03100v1 Announce Type: cross Abstract: Modern web applications increasingly expose accessibility barriers through interaction flows rather than static page snapshots. Keyboard traps, focus

Governed Caste Reassignment in Heterogeneous Swarms: An Asymmetric-Trust Protocol with Audited Operator Countersignature

AgentsDGX agent

arXiv:2607.04634v1 Announce Type: new Abstract: In heterogeneous robot swarms, caste reassignment (rebinding a robot to a new capability-bound role) is a high-frequency runtime event driven by battery

Heaviside Continuity of Rolling Coefficients for Eliminating Epistemic Entropy in Large Language Models

AgentsDGX agent

arXiv:2607.04562v1 Announce Type: new Abstract: Large language models (LLMs) generate fluent outputs that can be wrong. Unlike humans, who often exhibit cues when providing false information, LLMs pro

Multi-Large Language Model Orchestrated Severity Assessment of Clinical Records (MOSAIC)

Local AiDGX agent

arXiv:2607.05032v1 Announce Type: new Abstract: Background: Disease severity is a multidimensional construct difficult to capture with rule-based approaches in Electronic Healthcare Records (EHR). Age

PromptPET: Privacy-Utility Optimized Prompt Obfuscation

AgentsDGX agent

arXiv:2607.02932v1 Announce Type: cross Abstract: Privacy is an important challenge when users interact with AI chatbots, since users may share sensitive information, explicitly or implicitly, and AI

SMOCS: A Streaming Framework for Simplified Deployment, Monitoring, and Optimization of ML Systems in Production

AgentsDGX agent

arXiv:2607.02731v1 Announce Type: cross Abstract: Machine learning has demonstrated significant potential for real-time monitoring, optimization, and control of scientific facilities. However, deployi

TGRIP: A Text-Guided Approach to Vehicle Instance Prediction in Autonomous Driving

AgentsDGX agent

arXiv:2607.04812v1 Announce Type: new Abstract: Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the erro

This is the great thing about OpenWiki! It’ll maintain the docs for you automatically. Just setup a GitHub to run: openwiki —update and it’l…

AgentsDGX agent

This is the great thing about OpenWiki! It’ll maintain the docs for you automatically. Just setup a GitHub to run: openwiki —update and it’ll put up a PR once a day with changes / additions to your me

When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions

Model ReleasesDGX agent

arXiv:2607.03386v1 Announce Type: new Abstract: Agentic AI systems are increasingly used to edit, refine, and repair decision policies, but evaluating these edits is difficult when per-state expert ac

6 Jul 2026

big things happened over the weekend!

AgentsDGX agent

big things happened over the weekend! OpenWiki is at 1.7k stars in just 3 days! Right now it's just for codebases, but we're working to expand it to everything for memory. What do you want to see in a

How do you trace one number in a 200 page ESG report back to the exact page it came from? We dug into that with the @llama_index team behind…

AgentsDGX agent

How do you trace one number in a 200 page ESG report back to the exact page it came from? We dug into that with the @llama_index team behind LiteParse. We tested five ways to retrieve evidence across

I predict 50% of companies will need new leadership, because the old management style won't work in the era of AI. That's exactly why I'm se…

AgentsDGX agent

I predict 50% of companies will need new leadership, because the old management style won't work in the era of AI. That's exactly why I'm seeing 95%+ of so-called 'AI Transformation' initiatives fail.

Run MiniMax models on Amazon Bedrock

AgentsDGX agent

In this post, we walk through how to get started with MiniMax models on Amazon Bedrock, including the capabilities supported by these models, the service tiers available, how on-demand inference scale

3 Jul 2026

.@aiDotEngineer World's Fair was one of the most unique, interesting conferences I've been to: - incredible conversations with builders - hi…

AgentsDGX agent

.@aiDotEngineer World's Fair was one of the most unique, interesting conferences I've been to: - incredible conversations with builders - hilarious & creative touches (shout out @swyx) from a flash mo

Autonomous discovery of traffic laws with AI traffic scientists

AgentsDGX agent

arXiv:2607.01639v1 Announce Type: new Abstract: Universal traffic laws describe recurrent patterns in congestion, mobility and driving behavior across cities, providing a scientific basis for transpor

Beyond Supervised Clarification: Input Rewriting with LLMs for Dialogue Discourse Parsing

AgentsDGX agent

arXiv:2607.01964v1 Announce Type: new Abstract: Rewriting inputs to improve frozen downstream models has become a common strategy in modern NLP pipelines. Prior work on incremental dialogue discourse

Certified World Models as Sensing Clocks: Drift-Aware Deadlines for Active Perception

AgentsDGX agent

arXiv:2607.01537v1 Announce Type: new Abstract: Certified world models estimate how long their predictions remain valid. We turn this validity horizon into an operational sensing clock: a rule for whe

CommonRoad-Game: A Human-in-the-Loop Simulation Framework for Autonomous Driving

AgentsDGX agent

arXiv:2607.01382v1 Announce Type: new Abstract: Motion planning algorithms should be evaluated in human-in-the-loop environments to ensure they produce safe and efficient behaviors during interactions

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair

Model ReleasesDGX agent

arXiv:2607.01916v1 Announce Type: new Abstract: Large language model agents can repair real repository issues, but they often spend large context budgets on whole-file reads, broad searches, and long

Grounded Optimization: A Layered Engineering Framework for Reducing LLM Hallucination in Automated Personal Document Rewriting

AgentsDGX agent

arXiv:2607.01457v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly applied to resume optimization for applicant tracking systems, introducing hallucination failures distin

Language Models as Measurement Apparatus for Culture

AgentsDGX agent

arXiv:2607.02459v1 Announce Type: new Abstract: Language models are increasingly used to quantify cultural phenomena, but what makes such measurement distinctively cultural? This paper argues that NLP

Ophiuchus: Incentivizing Tool-augmented 'Think with Images' for Joint Medical Segmentation, Understanding and Reasoning

AgentsDGX agent

arXiv:2512.14157v2 Announce Type: replace Abstract: Recent medical MLLMs have made significant progress in generating step-by-step textual reasoning chains. However, they still struggle with complex c

OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration

Model ReleasesDGX agent

arXiv:2607.01531v1 Announce Type: new Abstract: Learning how an environment behaves from interaction is central to building agents that adapt to unfamiliar tasks. World models learned with deep networ

Prompt engineering is costing you money. Learn how to fine-tune your models. Stuffing a lot of text into every single API call slows down yo…

AgentsDGX agent

Prompt engineering is costing you money. Learn how to fine-tune your models. Stuffing a lot of text into every single API call slows down your app because the model has to process all those tokens bef

Recursive Models for Long-Horizon Reasoning

AgentsDGX agent

arXiv:2603.02112v2 Announce Type: replace-cross Abstract: Modern language models reason within bounded context, an inherent constraint that poses a fundamental barrier to long-horizon reasoning. We id

TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution

Model ReleasesDGX agent

arXiv:2607.02469v1 Announce Type: cross Abstract: Software tests and code evolve together: a code change should be followed by new or updated tests that record the new software behavior. Yet existing

2 Jul 2026

3 years ago I gave a talk at the first @aiDotEngineer conference on 'Advanced RAG' techniques in order to work around the limitations of nai…

Model ReleasesDGX agent

3 years ago I gave a talk at the first @aiDotEngineer conference on 'Advanced RAG' techniques in order to work around the limitations of naive RAG. It's insane how much the world has changed since the

Active Sensing for RIS-Aided Tracking and Power Control: A Hybrid Neuroevolution and Supervised Learning Approach

AgentsDGX agent

arXiv:2607.00056v1 Announce Type: cross Abstract: This paper studies energy efficient tracking of power-limited mobile users with the assistance of a Reconfigurable Intelligent Surface (RIS). Since lo

Best practices for multi-turn reinforcement learning in Amazon SageMaker AI

AgentsDGX agent

In this post, we share best practices for reliable multi-turn RL training. We cover how to build a training environment you can trust, set up an external evaluation, design a reward aligned with the e

Building on Devin's new capabilities, today we're announcing the Security Vulnerability Remediation Program: a tailored engagement to take y…

AgentsDGX agent

Building on Devin's new capabilities, today we're announcing the Security Vulnerability Remediation Program: a tailored engagement to take your security backlog towards zero in six weeks. Remediation

Last night we hosted the BabyAGI x Physical AI Happy Hour in SF with @yoheinakajima

AgentsDGX agent

Yohei Nakajima hosted a networking event in San Francisco bringing together the BabyAGI and Physical AI communities for casual conversation and relationship-building. The event likely focused on discu

Technical Report: Asynchronous Distributed Trajectory Estimation of Multi-Robot Systems

ResearchDGX agent

arXiv:2607.01106v1 Announce Type: new Abstract: Distributed trajectory estimation arises in many applications across robotics, but existing implementations typically do not consider asynchrony in agen

VLM-AR3L: Vision-Language Models for Absolute and Relative Rewards in Reinforcement Learning

AgentsDGX agent

arXiv:2607.00483v1 Announce Type: new Abstract: Designing effective reward functions remains a major challenge in reinforcement learning (RL), particularly in open-ended environments where task goals

1 Jul 2026

A Technical Typology of AI Systems in Public Administration

AgentsDGX agent

arXiv:2606.31755v1 Announce Type: cross Abstract: Research on artificial intelligence (AI) in the public sector often treats 'AI' as a single category, neglecting technical distinctions between differ

An AI-Based Solution for Secure Service Provisioning in IoT

AgentsDGX agent

arXiv:2606.30701v1 Announce Type: cross Abstract: As the Internet of Things (IoT) continues its rapid expansion, the attack surface grows accordingly, with emerging threats targeting smart objects and

Ask the World Before Acting: Budgeted Environment Probing for World-Model Calibration

SafetyDGX agent

arXiv:2606.31422v1 Announce Type: new Abstract: Long-horizon language agents do not only choose actions; they carry a private model of the world from one decision to the next. When that model drifts,

Compositional Concept-Based Neuron-Level Interpretability for Deep Reinforcement Learning

AgentsDGX agent

arXiv:2502.00684v2 Announce Type: replace-cross Abstract: Deep reinforcement learning (DRL) has successfully addressed many complex control problems. However, the neural networks representing policies

Continuous-Space Roadmap Generation for Mobile Robot Fleets with Distance Constraints and Geometry-Aware Discretization

AgentsDGX agent

arXiv:2511.07175v2 Announce Type: replace Abstract: Efficient routing of mobile robot fleets requires roadmaps with high redundancy, short path lengths, and sufficient node and edge clearance for conf

CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization

Model ReleasesDGX agent

arXiv:2606.31219v1 Announce Type: new Abstract: Cellular vehicle-to-everything (C-V2X) enables cooperative perception, prediction, and planning beyond the field of view of individual agents. However,

Long-term Traffic Simulation via Structured Autoregressive Modeling

SafetyDGX agent

arXiv:2606.31209v1 Announce Type: new Abstract: Interactive traffic simulation is a vital world model for autonomous driving. A central challenge in long-horizon simulation is modeling sustained multi

Motion Planning in Compressed Representation Spaces

AgentsDGX agent

arXiv:2606.30940v1 Announce Type: cross Abstract: Deep learning methods have vastly expanded the capabilities of motion planning in robotics applications, as learning priors from large-scale data has

MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning

Model ReleasesDGX agent

arXiv:2606.31073v1 Announce Type: new Abstract: Large language models (LLMs) provide a promising interface for high-level robotic task planning, but their use in multi-UAV collaboration remains diffic

SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QA

AgentsDGX agent

arXiv:2504.07385v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) become increasingly used for question-answering (QA), relying on static, pre-annotated references for evaluati

Scalable Behaviour Cloning on Browser Using via Skill Distillation

ApplicationsDGX agent

arXiv:2606.32014v1 Announce Type: new Abstract: Internet users collectively perform an enormous range of skilled work through web browsers, from software development and document editing to search, fo

The Harness should be Model agnostic. The Knowledge it builds around you is your moat. You should not have to migrate or give that up just b…

AgentsDGX agent

The Harness should be Model agnostic. The Knowledge it builds around you is your moat. You should not have to migrate or give that up just because a model is banned or just not good. Right now in my R

The Sensorimotor World Model (https://arxiv.org/abs/2606.20104): a deep dive into the role of inverse dynamics modeling as an anti-collapse …

AgentsDGX agent

The Sensorimotor World Model (https://arxiv.org/abs/2606.20104): a deep dive into the role of inverse dynamics modeling as an anti-collapse regularization for JEPAs. IDM is weaker than SIGReg as it do

This is pretty concerning. You could still do this at the API level to some degree, but they seemingly just blatantly put it right into the …

Model ReleasesDGX agent

This is pretty concerning. You could still do this at the API level to some degree, but they seemingly just blatantly put it right into the code? This is why open harnesses and agents are a much bette

This was an awesome collab between 7 companies 🔥 Huge shoutout to @browserbase @modal @braintrust @turbopuffer @p0 @cursor_ai and of course…

AgentsDGX agent

This was an awesome collab between 7 companies 🔥 Huge shoutout to @browserbase @modal @braintrust @turbopuffer @p0 @cursor_ai and of course our very own @llama_index teams. For those who attended - ho

Unveiling Transferability in Trajectory Prediction via Latent Scene Embeddings

AgentsDGX agent

arXiv:2606.30777v1 Announce Type: new Abstract: The growing availability of trajectory datasets has fueled major advances in data-driven motion prediction. Yet, models trained on one dataset often fai

Visual Prompt Discovery via Semantic Exploration

AgentsDGX agent

arXiv:2603.16250v2 Announce Type: replace-cross Abstract: LVLMs encounter significant challenges in image understanding and visual reasoning, leading to critical perception failures. Visual prompts, w

we built the @hyperspell filesystem around this principle every source in your company feeds into a context graph where we resolve conflicts…

AgentsDGX agent

we built the @hyperspell filesystem around this principle every source in your company feeds into a context graph where we resolve conflicts and assign permissions there’s a filesystem on top of the g

We call ours Operational Language Wiki, essentially a pattern for building memory to fill in the delta between the “dialect” of a team vs th…

AgentsDGX agent

We call ours Operational Language Wiki, essentially a pattern for building memory to fill in the delta between the “dialect” of a team vs the “world language” that the LLM is trained on. Use it to cap

What is your definition of a forward deployed engineer? asks @latentspacepod. @natalie_meurer: 'That is really the point of my session: the …

AgentsDGX agent

What is your definition of a forward deployed engineer? asks @latentspacepod. @natalie_meurer: 'That is really the point of my session: the role lacks a consistent definition. If you look at its histo

World-Model Collapse as a Phase Transition

Model ReleasesDGX agent

arXiv:2606.31399v1 Announce Type: new Abstract: Water looks unchanged as it warms, then at a critical point it boils. We ask whether long-horizon language agents show an analogous transition in their

← Previous
1…184185186187188…300
Next →