AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
9 Apr 2026

The wins came from three places: 1. Pre-computing upfront in a clever way that doesn't block the main thread 2. Avoiding NAPI overhead for s…

Model ReleasesDGX agent

The wins came from three places: 1. Pre-computing upfront in a clever way that doesn't block the main thread 2. Avoiding NAPI overhead for small result sets 3. Iteratively asking Claude to find perf i

8 Apr 2026

Hands on, concrete guide (with code!) for harness hill climbing with evals

TutorialsDGX agent

LangChain's 'Better Harness' tutorial, authored by Product Manager Vivek Trivedy and shared by Harrison Chase (@hwchase17), presents a hands-on, code-driven guide for using evaluations (evals) as a...

14 Aug 2026

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
A fantastic show as always, thanks for having me on @PeterMcCormack !
TutorialsDGX agent

A fantastic show as always, thanks for having me on @PeterMcCormack ! AI Has Escaped... - OpenAI agents escaped from their sandbox - They sent messages to other agents on how to escape - They hacked a

Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry

Model ReleasesDGX agent

arXiv:2608.12753v1 Announce Type: new Abstract: We study decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmet

FlowLOB: Efficient and Controllable Limit Order Book Generation with Flow Matching

AgentsDGX agent

arXiv:2608.13096v1 Announce Type: new Abstract: Limit order book (LOB) simulators are most useful to practitioners when they combine realistic market dynamics, computationally efficient sampling, cont

I’ve tried Cursor, Claude Code, Ollama, etc. — but I still don’t know how to use AI effectively for coding

Model ReleasesDGX agent

I’ve experimented with Cursor, Antigravity, Claude Code/CLI, Ollama, Gemma 4B, MiniMax, Hermes Agent, and different local/cloud models. My problem isn’t knowing what these tools are—I don't know how t

Mind the Context: Continual Learning of Socially Appropriate Robot Actions via Environmental-Social Disentanglement

AgentsDGX agent

arXiv:2608.13448v1 Announce Type: new Abstract: Social robots are expected to operate across diverse environments, where similar arrangements can imply different socially appropriate actions, e.g., st

PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research

Model ReleasesDGX agent

arXiv:2512.19799v2 Announce Type: replace Abstract: Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging because

13 Aug 2026

Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier

AgentsDGX agent

arXiv:2608.11247v1 Announce Type: new Abstract: Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improvi

D3D-GEN: Robot-Aware Domain-Grounded Interactive 3D World Generation for Social Robotics

AgentsDGX agent

arXiv:2608.11876v1 Announce Type: new Abstract: Training and validation of Embodied AI for social navigation critically depends on realistic simulation environments, yet many current approaches fail t

DaViNCi: A Dataset Towards Outdoor Vision-and-Language Navigation with Continuous Actions and Dynamic Elements

AgentsDGX agent

arXiv:2608.11901v1 Announce Type: new Abstract: Vision-and-Language Navigation (VLN) has progressively expanded from indoor to outdoor environments. However, existing outdoor VLN datasets still rely o

— Google AI Pro and Ultra subscribers can experience 3.7 Flash today via Spark in the @GeminiApp — Access the model in the Gemini Enterprise…

Model ReleasesDGX agent

— Google AI Pro and Ultra subscribers can experience 3.7 Flash today via Spark in the @GeminiApp — Access the model in the Gemini Enterprise Agent Platform and Gemini Enterprise app — Build in the Gem

LLMs in Process Diagram Engineering: From Optimal PFDs to Validated P&IDs

AgentsDGX agent

arXiv:2608.11220v1 Announce Type: new Abstract: Nowadays, the creation of a process flow diagram (PFD) and its subsequent transformation into a piping and instrumentation diagram (P&ID) is predominant

Self-evolving network verifiers

AgentsDGX agent

arXiv:2608.11340v1 Announce Type: cross Abstract: Symbolic network verifiers can reason about correctness across vast spaces of routing inputs and failures, but only for the protocols and features an

XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication

Model ReleasesDGX agent

arXiv:2608.11676v1 Announce Type: new Abstract: Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by reducing redun

12 Aug 2026

Behavioral Inference at Scale: The Fundamental Asymmetry Between Motivations and Belief Systems

Model ReleasesDGX agent

arXiv:2509.05624v3 Announce Type: replace-cross Abstract: How much information about an agent's underlying values can be recovered from its observable behavior? This question matters for any approach

Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning

AgentsDGX agent

arXiv:2608.10438v1 Announce Type: new Abstract: Large language models increasingly rely on external tools to access up-to-date information, perform computation, and interact with the outside world. Fo

Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4

HardwareDGX agent

arXiv:2608.10103v1 Announce Type: cross Abstract: High-performance Tensor Core kernels rely on a low-level PTX pipeline built from asynchronous data movement with cp.async, warp-level matrix loads wit

Live now: our Memory & Continual Learning Track from AI Engineer World's Fair 2026. Thesis: we scaled intelligence and got the world's smart…

TutorialsDGX agent

Live now: our Memory & Continual Learning Track from AI Engineer World's Fair 2026. Thesis: we scaled intelligence and got the world's smartest novice. https://www.youtube.com/watch?v=iqloyWCGYQQ&list

Nebius shares jump 34% on continued AI infrastructure demand

AgentsDGX agent

Shares of Nebius Group NV closed 34% higher today after it reported second-quarter earnings that topped expectations across the board. The Netherlands-based company operates a cloud platform optimized

The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse

AgentsDGX agent

arXiv:2608.10186v1 Announce Type: cross Abstract: LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests l

11 Aug 2026

A Communication-Efficient Digital Twin Framework for PSO-Based Swarm Navigation and Obstacle Avoidance

AgentsDGX agent

arXiv:2406.19930v4 Announce Type: replace Abstract: Swarm-based target localization in industrial environments faces two major challenges: navigating obstacle-rich spaces and managing intensive commun

Benchmarking In-context Experiential Learning Through Repeated Product Recommendations

Model ReleasesDGX agent

arXiv:2511.22130v2 Announce Type: replace Abstract: To navigate ever-shifting real-world environments, agents must grapple with incomplete knowledge and adapt their strategies through experience. Howe

Had so much fun giving this talk at @sequoia about @harvey’s moneyball approach to building a research lab. The biggest mistake I made in th…

AgentsDGX agent

Had so much fun giving this talk at @sequoia about @harvey’s moneyball approach to building a research lab. The biggest mistake I made in the early days of Harvey was trying to play the Yankees baseba

IntelliAudit: Using Large Language Models to Evaluate Audit Controls

AgentsDGX agent

arXiv:2608.07688v1 Announce Type: new Abstract: IT audits require auditors to judge whether heterogeneous organizational evidence satisfies semantic security and compliance controls. This judgment is

Skills in Weights, Memory in Code: Hybrid Learning for Memory-Dependent Robot Manipulation

AgentsDGX agent

arXiv:2608.09410v1 Announce Type: new Abstract: Modern vision-language-action (VLA) policies have acquired broad manipulation skills, but typically generate each action chunk from the current observat

The Capability Ladder: A Curriculum-Modernization Framework for Workforce Readiness in the AI Era

AgentsDGX agent

arXiv:2608.07779v1 Announce Type: new Abstract: Artificial intelligence is changing the task composition of computing work faster than curricula and training typically adapt. This is a curriculum-fram

TRACE: TRajectory Attribution for Automated Context Engineering

Model ReleasesDGX agent

arXiv:2608.09153v1 Announce Type: new Abstract: Production AI agents fail when their context sources -- system prompts, knowledge bases, tool descriptions, and procedural skills -- contain errors or g

VDGR-RAG: Vectors, Directories, Graphs, and Reflection Are All You Need for Unified Reasoning over Hierarchical Enterprise Knowledge

AgentsDGX agent

arXiv:2608.07994v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is essential for enterprise knowledge question answering (QA), particularly in domains with complex product documen

10 Aug 2026

Interaction Creates Dynamical AI Behavior Absent in Isolation

ResearchDGX agent

arXiv:2608.07457v1 Announce Type: new Abstract: What will happen when AI agents interact in daily life, e.g. when one AI starts bossing another around? We find a counterintuitive answer that opens new

Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving

AgentsDGX agent

arXiv:2603.06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their r

Vehicle routing problem using deep reinforcement learning - A case study about truck planning in the industry

AgentsDGX agent

arXiv:2608.06668v1 Announce Type: new Abstract: As an important component of the supply chain industry, transportation has experienced rapid development in the past decade with the assistance of digit

8 Aug 2026

Auto mode is now the default in Claude Code for Pro, Max, and Team plans

Model ReleasesDGX agent

Auto mode is now the default in Claude Code for Pro, Max, and Team plans Anthropic are really confident in Claude Code's auto mode, to the point that they are making it the default setting for new ses

7 Aug 2026

Computationally Efficient Collaborative Communication Via Regularity-Based Coarsening

AgentsDGX agent

arXiv:2608.05327v1 Announce Type: cross Abstract: Our results show that the existence of a short high-utility protocol already suffices for efficient communication. In particular, in a game with n pos

How Cohere Health digitizes clinical policies using Amazon Bedrock AgentCore

SafetyDGX agent

In this post, you learn how Cohere Health built a multi-tenant agentic architecture on AgentCore using AgentCore Runtime’s secure MicroVM isolation, unified tool access through AgentCore Gateway, Agen

Recursive Synthesis for Long-Horizon Terminal Tasks

Model ReleasesDGX agent

arXiv:2608.05466v1 Announce Type: new Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because ea

upgraded my stack, and i can now work on almost anything from anywhere hands free: - talk to chief of staff (via remote codex voice or text)…

Model ReleasesDGX agent

upgraded my stack, and i can now work on almost anything from anywhere hands free: - talk to chief of staff (via remote codex voice or text) - chief assigns tasks to managers of various projects - man

Zero-code, low-cost data ingestion: New BigQuery DTS capabilities

AgentsDGX agent

In a fast-paced digital economy, data is your most critical engine. Yet, many enterprises find themselves trapped in a costly paradox, spending over 100 hours a week building and fixing fragile, in-ho

6 Aug 2026

i guess this is a good time to mention that smol forge is open for the first 100 alpha users. get your usernames! (tire kickers who dont mak…

AgentsDGX agent

i guess this is a good time to mention that smol forge is open for the first 100 alpha users. get your usernames! (tire kickers who dont make any commits will be kicked out by eod) point clanker to fo

Improving Auto-Design of Neural PDE Solvers with a Domain-Specific Language

AgentsDGX agent

arXiv:2608.04384v1 Announce Type: new Abstract: Neural PDE solver auto-design is fundamentally a search-space representation problem. In the space of unrestricted Python programs, valid solvers form a

OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling

AgentsDGX agent

arXiv:2608.05141v1 Announce Type: new Abstract: Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon

When Shared Rollouts Fail in Defensive Driving Evaluation: A NAVSIM Score Basis Audit

AgentsDGX agent

arXiv:2608.04896v1 Announce Type: new Abstract: Defensive driving scores are useful only when they preserve distinctions between policies that observe surrounding actors and those that do not. Re-simu

5 Aug 2026

EFX Allocation In (Multi)Hypergraphs

ResearchDGX agent

arXiv:2608.03171v1 Announce Type: cross Abstract: We study fair allocations of indivisible goods among agents with heterogeneous monotone valuations. As fair we consider the allocations that are envy-

From Routes to Steps: Separating Semantic Progress from Local Execution in Vision-and-Language Navigation

SafetyDGX agent

arXiv:2608.03143v1 Announce Type: new Abstract: Vision-and-Language Navigation (VLN) requires an agent to follow a route-level instruction by executing its constituent steps from egocentric visual obs

Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls

Model ReleasesDGX agent

arXiv:2608.03071v1 Announce Type: new Abstract: Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting the right tool

'OCR is just a feature now. Frontier models will eat it.' We hear this constantly. The data says otherwise. Across three GPT generations, pa…

AgentsDGX agent

'OCR is just a feature now. Frontier models will eat it.' We hear this constantly. The data says otherwise. Across three GPT generations, parsing accuracy gained ~24 points, while cost per page 4x'd.

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks

AgentsDGX agent

arXiv:2607.28587v2 Announce Type: replace-cross Abstract: SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability. They typically follow a common construction pipeli

Residual Flow Matching with Dynamic Cross-Interaction for 3D Multi-Person Motion Prediction

AgentsDGX agent

arXiv:2608.03379v1 Announce Type: new Abstract: 3D multi-person motion prediction requires modeling both individual kinematics and inter-person interactions. While Flow Matching is effective for multi

4 Aug 2026

AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their Impact on Automated Assessment

SafetyDGX agent

arXiv:2608.00717v1 Announce Type: cross Abstract: Rubric-based AI systems for thesis assessment use criterion weights to assign different levels of importance to evaluation criteria. These weights are

Don't Offer What Can't Be Done: Deterministic Executability Gating for LLM Skill Selection at Scale

ApplicationsDGX agent

arXiv:2608.01050v1 Announce Type: cross Abstract: Production LLM agents that select from large skill libraries face a limitation that semantic relevance alone cannot resolve: a skill may match a user'

Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races

Model ReleasesDGX agent

arXiv:2608.01193v1 Announce Type: cross Abstract: An AI development race creates a multi-agent safety dilemma. Each company can develop slowly and safely, or move faster while taking a risk that may r

Inference-Time Policy Alignment for Fair Reinforcement Learning

SafetyDGX agent

arXiv:2608.00175v1 Announce Type: new Abstract: Deep reinforcement learning (RL) agents achieve strong performance by optimizing scalar reward functions. However, once deployed, the policies of these

PB^2: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning

AgentsDGX agent

arXiv:2506.13741v2 Announce Type: replace-cross Abstract: Preference-based reinforcement learning (PbRL) has emerged as a promising approach for learning behaviors from human feedback without predefin

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise

Model ReleasesDGX agent

arXiv:2608.02372v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in task-oriented dialogue systems that support multi-step decision-making in high-stakes domains

RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Review

AgentsDGX agent

arXiv:2608.00005v1 Announce Type: new Abstract: Peer review at major venues is under unprecedented submission pressure, motivating the use of large language models (LLMs) as review assistants. Existin

We were proud to host “Build with Frontier Intelligence,” http://Z.ai ’s first community meetup, together with @AISingapore, at Tencent’s ve…

AgentsDGX agent

We were proud to host “Build with Frontier Intelligence,” http://Z.ai ’s first community meetup, together with @AISingapore, at Tencent’s venue. During the event, http://Z.ai shared insights into the

3 Aug 2026

Data Turnstile: A Scalable Open Framework for Function-Calling Data Generation

Model ReleasesDGX agent

arXiv:2607.29250v1 Announce Type: new Abstract: Small language models (SLMs) are attractive for agentic deployment due to low latency, reduced cost, and on-device privacy, yet they struggle with tool-

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat

Model ReleasesDGX agent

arXiv:2607.29577v1 Announce Type: new Abstract: Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich tactical reas

UltraSAM3: A Concept-Driven Foundation Model for Universal Ultrasound Image Segmentation

AgentsDGX agent

arXiv:2607.29200v1 Announce Type: new Abstract: Ultrasound imaging has become increasingly widespread in clinical practice due to its portability, low cost and real-time capability, making ultrasound

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

SafetyDGX agent

arXiv:2607.28678v1 Announce Type: new Abstract: Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally

← Previous
1…173174175176177…300
Next →