AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Model Releases

Prompt reinforcing for long-term planning of large language models

DGX agent

arXiv:2510.05921v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved remarkable success in a wide range of natural language processing tasks and can be adapted through prompt

model-releasesarxiv-cs-cl
10 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Safe Large-Scale Robust Nonlinear MPC in Milliseconds via Reachability-Constrained System Level Synthesis on the GPU

DGX agent

arXiv:2604.07644v1 Announce Type: new Abstract: We present GPU-SLS, a GPU-parallelized framework for safe, robust nonlinear model predictive control (MPC) that scales to high-dimensional uncertain rob

safetyarxiv-cs-ro
10 Apr 2026
Research

Say Something Else: Rethinking Contextual Privacy as Information Sufficiency

DGX agent

arXiv:2604.06409v1 Announce Type: cross Abstract: LLM agents increasingly draft messages on behalf of users, yet users routinely overshare sensitive information and disagree on what counts as private.

researcharxiv-cs-ai
10 Apr 2026
Model Releases

Spatio-Temporal Grounding of Large Language Models from Perception Streams

DGX agent

arXiv:2604.07592v1 Announce Type: new Abstract: Embodied-AI agents must reason about how objects move and interact in 3-D space over time, yet existing smaller frontier Large Language Models (LLMs) st

model-releasesarxiv-cs-ro
10 Apr 2026
Safety

Temporal Inversion for Learning Interval Change in Chest X-Rays

DGX agent

arXiv:2604.04563v2 Announce Type: replace-cross Abstract: Recent advances in vision--language pretraining have enabled strong medical foundation models, yet most analyze radiographs in isolation, over

safetyarxiv-cs-ai
10 Apr 2026
Safety

Towards Privacy-Preserving Large Language Model: Text-free Inference Through Alignment and Adaptation

DGX agent

arXiv:2604.06831v1 Announce Type: cross Abstract: Current LLM-based services typically require users to submit raw text regardless of its sensitivity. While intuitive, such practice introduces substan

safetyarxiv-cs-ai
10 Apr 2026
Safety

Towards provable probabilistic safety for scalable embodied AI systems

DGX agent

arXiv:2506.05171v3 Announce Type: replace-cross Abstract: Embodied AI systems, comprising AI models and physical plants, are increasingly prevalent across various applications. Due to the rarity of sy

safetyarxiv-cs-ai
10 Apr 2026
Safety

VLMShield: Efficient and Robust Defense of Vision-Language Models against Malicious Prompts

DGX agent

arXiv:2604.06502v1 Announce Type: new Abstract: Vision-Language Models (VLMs) face significant safety vulnerabilities from malicious prompt attacks due to weakened alignment during visual integration.

safetyarxiv-cs-lg
10 Apr 2026
Model Releases

Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework

DGX agent

arXiv:2608.11891v1 Announce Type: cross Abstract: Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual computing. Asse

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams

DGX agent

arXiv:2608.12262v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a fre

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration

DGX agent

arXiv:2608.11210v1 Announce Type: new Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers

model-releasesarxiv-cs-ai
13 Aug 2026
Safety

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

DGX agent

arXiv:2608.12063v1 Announce Type: cross Abstract: Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely

safetyarxiv-cs-ai
13 Aug 2026
Model Releases

OEIS Open: How many conjectures can language models turn into theorems?

DGX agent

arXiv:2608.11941v1 Announce Type: new Abstract: We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al. Whereas these conje

model-releasesarxiv-cs-ai
13 Aug 2026
Safety

Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits

DGX agent

arXiv:2608.11410v1 Announce Type: new Abstract: Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Err

safetyarxiv-cs-lg
13 Aug 2026
Model Releases

A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language

DGX agent

arXiv:2608.10588v1 Announce Type: cross Abstract: Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

An adaptive and evolvable deep reinforcement learning framework for weather prediction

DGX agent

arXiv:2608.09948v1 Announce Type: cross Abstract: No single AI weather model excels at all variables, pressure levels, and lead times. Rather than building yet another architecture, we reframe the for

model-releasesarxiv-cs-lg
12 Aug 2026
Model Releases

Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxy

DGX agent

arXiv:2608.10532v1 Announce Type: cross Abstract: Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server

model-releasesarxiv-cs-lg
12 Aug 2026
Model Releases

DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation

DGX agent

arXiv:2608.10636v1 Announce Type: cross Abstract: Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Pri

model-releasesarxiv-cs-cl
12 Aug 2026
Safety

Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving

DGX agent

arXiv:2608.10386v1 Announce Type: new Abstract: Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias. While world mod

safetyarxiv-cs-lg
12 Aug 2026
Safety

ELMER: Evolutionary Language Model that Explores and Refines

DGX agent

arXiv:2608.10196v1 Announce Type: cross Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is a

safetyarxiv-cs-ai
12 Aug 2026
Model Releases

ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering

DGX agent

arXiv:2608.10679v1 Announce Type: cross Abstract: Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, however, are wor

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

FormStruct-Bench:A Hierarchical and Diagnostic Benchmark for Table-Form Document Structure Recognition

DGX agent

arXiv:2608.10396v1 Announce Type: new Abstract: Transforming table-form documents into machine-processable records requires recovering not only their visible content but also the multilevel structure

model-releasesarxiv-cs-cv
12 Aug 2026
Safety

IO Factory: Simulating AI-Enabled Influence Campaigns at Scale

DGX agent

arXiv:2608.10920v1 Announce Type: new Abstract: We introduce IO Factory, an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes. The threat

safetyarxiv-cs-ai
12 Aug 2026
Model Releases

Rescene: band-limited stochastic forcing turns a frozen neural weather operator into a climate emulator

DGX agent

arXiv:2608.09971v1 Announce Type: cross Abstract: Over the past few years, the rapid development of machine learning (ML) models for weather forecasting has produced deterministic models whose medium-

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom

DGX agent

arXiv:2608.10145v1 Announce Type: new Abstract: LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached o

model-releasesarxiv-cs-lg
12 Aug 2026
Model Releases

V-FiLLM: Verified Financial LLM Reasoning Benchmark

DGX agent

arXiv:2608.11047v1 Announce Type: new Abstract: While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains compar

model-releasesarxiv-cs-ai
12 Aug 2026
Safety

Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

DGX agent

arXiv:2608.10703v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream dec

safetyarxiv-cs-ai
12 Aug 2026
Model Releases

ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

DGX agent

arXiv:2607.10180v2 Announce Type: replace-cross Abstract: We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception. T

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization

DGX agent

arXiv:2608.07557v1 Announce Type: cross Abstract: Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent minimalist en

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Back to the Future: A workbook time machine for spread sheet creation benchmarks

DGX agent

arXiv:2608.07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to create derived obj

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models

DGX agent

arXiv:2608.08061v1 Announce Type: new Abstract: The key question in moral judgement is not simply whether someone chooses the 'right' answer, but how they decide what matters most when moral principle

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility

DGX agent

arXiv:2608.08691v1 Announce Type: new Abstract: Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but physical flexibility becomes dependable capacity

model-releasesarxiv-cs-ai
11 Aug 2026
Local Ai

HarnessWAM: Bridging Prediction and Deliberation in World Action Models

DGX agent

arXiv:2608.09516v1 Announce Type: new Abstract: World Action Models (WAMs) jointly learn environmental dynamics and robot actions, introducing priors over physical evolution into embodied control. How

local-aiarxiv-cs-ro
11 Aug 2026
Safety

Illusion of Alignment: Detecting Hidden Disagreement in Collaborative Dialogue

DGX agent

arXiv:2608.08210v1 Announce Type: new Abstract: Collaborative dialogue can end with apparent agreement while participants still differ on goals, assumptions, or execution plans, creating an extbf{illu

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

LLM-Guided Heuristic Design from Simulation Traces: A Case Study in Dynamic Production and AGV Scheduling

DGX agent

arXiv:2608.09343v1 Announce Type: new Abstract: Simulation-based optimization (SBO) evaluates executable policies under stochastic dynamics, but most methods treat the simulator as a black box: aggreg

model-releasesarxiv-cs-ai
11 Aug 2026
Safety

ML-Based Hierarchical Prediction for Practical Energy Scheduling in Dynamic NTN-WPT Systems

DGX agent

arXiv:2608.08804v1 Announce Type: cross Abstract: With advancements in long-distance wireless power transfer (WPT) and space-based energy technologies, integrating WPT into non-terrestrial networks (N

safetyarxiv-cs-lg
11 Aug 2026
Safety

The Anatomy of a Prompt Injection: A Component Model for Structured Analysis

DGX agent

arXiv:2608.07808v1 Announce Type: cross Abstract: Four years after prompt injection was first identified in 2022, attacks are still predominantly documented as verbatim strings rather than structured

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

An Exploratory Evaluation of LLM-Assisted Rewriting of Moderate-Complexity Financial Sentences for DisCoCat-Based Sentiment Analysis

DGX agent

arXiv:2608.07439v1 Announce Type: new Abstract: Quantum natural language processing (QNLP) provides a grammar-aware framework for text modeling, and Distributional Compositional Categorical (DisCoCat)

model-releasesarxiv-cs-cl
10 Aug 2026
Model Releases

Beyond Text Matching: Towards Reference-Free Evaluation for Human-Oriented Binary Reverse Engineering

DGX agent

arXiv:2608.07038v1 Announce Type: cross Abstract: Human-Oriented Binary Reverse Engineering (HOBRE) aims to transform decompiled pseudocode into a more human-friendly representation, thereby reducing

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing

DGX agent

arXiv:2608.07424v1 Announce Type: new Abstract: Test-time scaling is often implemented by spending more compute along one axis: sampling more solutions, extending a chain of thought, or applying a str

model-releasesarxiv-cs-ai
10 Aug 2026
Research

DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding

DGX agent

arXiv:2608.07067v1 Announce Type: new Abstract: Long-document understanding requires locating sparse and heterogeneous evidence across hundreds of pages, yet existing systems remain limited by static

researcharxiv-cs-ai
10 Aug 2026
Model Releases

FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings

DGX agent

arXiv:2608.07400v1 Announce Type: new Abstract: Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible and even numerically correct answer can be gro

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks

DGX agent

arXiv:2608.07411v1 Announce Type: new Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into

model-releasesarxiv-cs-ai
10 Aug 2026
Local Ai

Human-Centered Explainable AI for TinyML Edge Devices: A Pareto-Based Selection Framework with LLM-Guided Design

DGX agent

arXiv:2608.07091v1 Announce Type: cross Abstract: Edge Artificial Intelligence (Edge AI) enables the deployment of AI models directly on local edge devices, while such deployments are subject to stric

local-aiarxiv-cs-ai
10 Aug 2026
Model Releases

LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering

DGX agent

arXiv:2608.07370v1 Announce Type: new Abstract: Scientific literature is increasingly used as a knowledge source for language models, retrieval-augmented generation systems, and research assistants, b

model-releasesarxiv-cs-cl
10 Aug 2026
Model Releases

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

DGX agent

arXiv:2608.06867v1 Announce Type: new Abstract: No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment.

model-releasesarxiv-cs-cl
10 Aug 2026
Safety

MaskFlow: Precise, Consistent and Seamless Regional Image Editing

DGX agent

arXiv:2608.06929v1 Announce Type: cross Abstract: Regional image editing has attracted considerable attention for its spatial controllability. Although instruction-based and mask-reference-based editi

safetyarxiv-cs-ai
10 Aug 2026
Safety

MolBioKG: Grounding Out-of-Graph Molecules in Biomedical Knowledge Graphs via Multi-Resolution Structural Anchoring

DGX agent

arXiv:2608.06713v1 Announce Type: new Abstract: Biomedical knowledge graphs (KGs) accelerate drug discovery, but standard pipelines assume query molecules already exist as graph entities, leaving unre

safetyarxiv-cs-ai
10 Aug 2026
← Previous
1…185186187188189…233
Next →