AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,770 results
Model Releases

Automated Report-Derived Oncology VQA Benchmark for Evaluating Vision-Language Models on 3D Medical Imaging

DGX agent

arXiv:2606.02809v1 Announce Type: new Abstract: Evaluating vision-language models (VLMs) on medical images requires benchmarks that are clinically grounded, scalable, and controlled for evaluation con

model-releasesarxiv-cs-cv
3 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models

DGX agent

arXiv:2606.03157v1 Announce Type: new Abstract: Large language models (LLMs) have been widely adopted in healthcare, yet they still encounter significant challenges in complex clinical decision-making

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

Discovering autonomous quantum error correction via deep reinforcement learning

DGX agent

arXiv:2511.12482v2 Announce Type: replace-cross Abstract: Quantum error correction is essential for fault-tolerant quantum computing. However, standard methods relying on active measurements may intro

safetyarxiv-cs-lg
3 Jun 2026
Safety

Human-in-the-Loop Contextual Bandits for Short-Term Rental Dynamic Pricing: Structural Equivalence of Historical Warm-Up and Approval-Gated Live Learning

DGX agent

arXiv:2606.02595v1 Announce Type: new Abstract: Dynamic pricing in short-term rental (STR) markets presents a distinctive challenge for online learning algorithms: pricing decisions carry significant

safetyarxiv-cs-lg
3 Jun 2026
Safety

Inference Cost Attacks for Retrieval-Augmented Large Language Models

DGX agent

arXiv:2606.02643v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG)-enhanced LLM systems, while powerful, introduce substantial inference costs due to the inclusion of an extra mult

safetyarxiv-cs-ai
3 Jun 2026
Model Releases

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing

DGX agent

arXiv:2606.02822v1 Announce Type: cross Abstract: Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-regis

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

AI-PROPELLER: Warehouse-Scale Interprocedural Code Layout Optimization with AlphaEvolve

DGX agent

arXiv:2606.00131v1 Announce Type: cross Abstract: Post-link optimizers (PLOs) such as Propeller and BOLT have demonstrated that precise, profile-guided code layout can extract significant performance

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

Beyond Text and Tables: Vision-Language Model Integration in ComProScanner for Extracting Materials Data from Scientific Figures with High Accuracy

DGX agent

arXiv:2606.00065v1 Announce Type: cross Abstract: Automated extraction of materials composition-property data from scientific literature has advanced considerably with the development of large languag

model-releasesarxiv-cs-ai
2 Jun 2026
Safety

ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic Education

DGX agent

arXiv:2512.05671v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have achieved remarkable success in dyadic (one-on-one) instruction, they face significant challenges in One-to-M

safetyarxiv-cs-cl
2 Jun 2026
Safety

Crazyflow: An Accurate, GPU-Accelerated, Differentiable Drone Simulator in JAX

DGX agent

arXiv:2606.01478v1 Announce Type: cross Abstract: High-quality, large-scale synthetic data from simulations is becoming a cornerstone for pushing the capabilities of robot algorithms. While aerial rob

safetyarxiv-cs-ai
2 Jun 2026
Safety

Explainable Data-driven Deep Reinforcement Learning Methods for Optimal Energy Management in Buildings

DGX agent

arXiv:2606.02049v1 Announce Type: new Abstract: The increasing integration of renewable energy sources into power systems, particularly in buildings equipped with photovoltaic (PV) panels and energy s

safetyarxiv-cs-ai
2 Jun 2026
Safety

From Graph Retrieval to Schema Realization: Counterfactual Validation for Text-to-SPARQL over Heterogeneous Knowledge Graphs

DGX agent

arXiv:2508.01815v2 Announce Type: replace-cross Abstract: Text-to-SPARQL maps natural-language questions to executable SPARQL queries over RDF knowledge graphs. While standard evaluations often fix th

safetyarxiv-cs-ai
2 Jun 2026
Safety

From 'Weak' Signals to Strong Models: Preference Delta Aggregation with LoRA Merging

DGX agent

arXiv:2606.00357v1 Announce Type: new Abstract: Training strong large language models (LLMs) requires high-quality supervision, which is often scarce. Recent work shows that paired preference data fro

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding

DGX agent

arXiv:2606.01019v1 Announce Type: cross Abstract: Large Language Model (LLM) generation remains expensive because autoregressive decoding calls the model once for each new token. Speculative decoding

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics

DGX agent

arXiv:2606.01502v1 Announce Type: cross Abstract: Frontier LLMs increasingly decide what a query attends to with a sparse-attention indexer that picks a few KV-cache blocks per query: attention's unit

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Reinforcement Learning for Optimal Experiment Design in Parameter Identification of Mechatronic Systems

DGX agent

arXiv:2606.00059v1 Announce Type: cross Abstract: Informative excitation signals are critical for accurate system identification of mechatronic systems, yet classical system identification (SI) approa

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Science Earth: Towards A Planet-Scale Operating System for AI-Native Scientific Discovery

DGX agent

arXiv:2606.01316v1 Announce Type: new Abstract: Scientific discovery demands intelligence, perseverance, and serendipity across vast search spaces. Today, top scientific capabilities remain siloed--on

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in co…

DGX agent

Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier. First is our text foundation mode

model-releasesfireworks-ai--x
2 Jun 2026
Tutorials

ToMAP: Training Opponent-Aware LLM Persuaders with Theory of Mind

DGX agent

arXiv:2505.22961v3 Announce Type: replace Abstract: Large language models (LLMs) have shown promising potential in persuasion, but existing works on training LLM persuaders are still preliminary. Nota

tutorialsarxiv-cs-cl
2 Jun 2026
Safety

TriAlign: Towards Universal Truth Consistency in Personalized LLM Alignment

DGX agent

arXiv:2606.01755v1 Announce Type: new Abstract: Personalized large language models adapt responses to users' preferences and social attributes, but can introduce substantial universal truth inconsiste

safetyarxiv-cs-ai
2 Jun 2026
Safety

World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications

DGX agent

arXiv:2606.00133v1 Announce Type: new Abstract: World models, internal simulators that learn the structure and dynamics of an environment, have emerged as a central paradigm in the pursuit of artifici

safetyarxiv-cs-lg
2 Jun 2026
Hardware

Five thoughts from Nvidia CEO Jensen Huang’s GTC Taipei 2026 keynote

DGX agent

Useful artificial intelligence has arrived, and if Nvidia Chief Executive Jensen Huang is right, it is about to reshape not only data centers but also the structure of the global economy and the tech

hardwaresiliconangle
1 Jun 2026
Model Releases

HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs

DGX agent

arXiv:2511.18760v2 Announce Type: replace Abstract: Informal mathematics has been central to modern large language model (LLM) reasoning, offering flexibility and efficient construction of arguments.

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Many research labs only consider inference efficiency after the fact. Step 3.7 Flash is a 196B MoE model, and built for inference from the s…

DGX agent

Many research labs only consider inference efficiency after the fact. Step 3.7 Flash is a 196B MoE model, and built for inference from the start by @StepFun_ai. Multi-Matrix Factorization Attention (M

model-releasesfireworks-ai--x
1 Jun 2026
Safety

Population-Free Pareto Tracking for Sample-Efficient Multi-Policy MORL

DGX agent

arXiv:2508.02217v2 Announce Type: replace Abstract: Multi-objective reinforcement learning (MORL) is a fundamental framework for real-world decision-making problems involving multiple conflicting crit

safetyarxiv-cs-lg
1 Jun 2026
Model Releases

How we contain Claude across products

DGX agent

How we contain Claude across products A complaint I often have about sandboxing products is that they are rarely thoroughly documented, and in the absence of detailed documentation it's hard to know h

model-releasessimon-willison
30 May 2026
Model Releases

Are LLMs Socially Adaptive? Contrasting Belief Evolution in Large Language Models and Humans

DGX agent

arXiv:2410.10398v3 Announce Type: replace-cross Abstract: As large language models (LLMs) increasingly engage in complex social interactions, ensuring that their behaviors align with human ethical pri

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials

DGX agent

arXiv:2510.04704v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promising potential in scientific research, enabling tasks ranging from knowledge retrieval to propert

model-releasesarxiv-cs-ai
29 May 2026
Safety

DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation

DGX agent

arXiv:2605.29522v1 Announce Type: new Abstract: As scientific literature grows rapidly, automated survey generation has become a key capability for AI scientists and human researchers. However, existi

safetyarxiv-cs-ai
29 May 2026
Local Ai

DynaGraph: Lightweight Multi-Model Interaction Framework via Dynamic Topological Reconfiguration

DGX agent

arXiv:2605.29511v1 Announce Type: cross Abstract: Tackling complex reasoning tasks typically relies on massive monolithic LLMs, which suffer from severe computational redundancy. While task decomposit

local-aiarxiv-cs-cl
29 May 2026
Applications

Enhancing Reinforcement Learning in 3D Environments through Semantic Segmentation: A Case Study in ViZDoom

DGX agent

arXiv:2511.11703v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) in 3D environments with high-dimensional sensory input poses two major challenges: (1) the high memory consumption

applicationsarxiv-cs-ai
29 May 2026
Model Releases

llama.cpp now has an official website: https://llama.app Our goal is to make local AI accessible to everyone, and improving the user experie…

DGX agent

llama.cpp now has an official website: https://llama.app Our goal is to make local AI accessible to everyone, and improving the user experience is a big part of that. On the new landing page you’ll fi

model-releasesclem-delangue--x
29 May 2026
Safety

Modularizing Educational LLM-Agency for Fostering Responsible Learning Assistance

DGX agent

arXiv:2605.30187v1 Announce Type: new Abstract: The widespread adoption of AI chatbots in education will drastically change learning, making responsible deployment a critical concern. While large lang

safetyarxiv-cs-ai
29 May 2026
Model Releases

RoboWits: Unexpected Challenges for Robotic Creative Problem Solving

DGX agent

arXiv:2605.30326v1 Announce Type: cross Abstract: The ability to reason, adapt, and creatively solve problems under unexpected challenges is essential for robots operating in real-world environments.

model-releasesarxiv-cs-ai
29 May 2026
Tutorials

Tip on Grok Build + Grok 4.3 VLM One of my critical tasks is to keep Grok VLM in the loop. Throwing a default system prompt usually yields p…

DGX agent

Tip on Grok Build + Grok 4.3 VLM One of my critical tasks is to keep Grok VLM in the loop. Throwing a default system prompt usually yields poor results due to lack of context. Here is how to scale: -

tutorialselon-musk--x
29 May 2026
Research

A New Era of Innovation: Google Research at I/O 2026

DGX agent

Google Research at I/O 2026 showcased new Gemini AI models including Gemini Omni, which can create content from any input starting with video, and Gemini 3.5 Flash, combining frontier intelligence wit

researchgoogle-research
28 May 2026
Safety

AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?

DGX agent

arXiv:2605.28255v1 Announce Type: new Abstract: AI systems are fallible, and humans can make mistakes in deciding whether to trust AI over their own judgment. Thus, improving human-AI collaboration re

safetyarxiv-cs-ai
28 May 2026
Model Releases

AlphaForgeBench: Benchmarking End-to-End Trading Strategy Design with Large Language Models

DGX agent

arXiv:2602.18481v2 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) has led to a surge of financial benchmarks, evolving from static knowledge evaluation to

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection

DGX agent

arXiv:2505.17654v4 Announce Type: replace-cross Abstract: E-commerce platforms increasingly rely on Large Language Models (LLMs) and Vision Language Models (VLMs) to detect illicit or misleading produ

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

ForestHG-Trace: Traceable Long-Horizon Ecological Reasoning over Large-Scale Forest Scenes

DGX agent

arXiv:2605.27590v1 Announce Type: new Abstract: Remote sensing question answering (RS-QA) often requires more than direct semantic prediction, especially in large-scale forest scenes where ecological

model-releasesarxiv-cs-cv
28 May 2026
Research

Human-AI Collaboration for Estimating Scientific Replicability

DGX agent

arXiv:2605.27394v1 Announce Type: cross Abstract: Determining whether published scientific findings can successfully be replicated is a long-standing challenge in the empirical sciences. Existing appr

researcharxiv-cs-ai
28 May 2026
Model Releases

OccuReward: LLM-Guided Occupant-Centric Reward Shaping for Demographic Equity in Grid-Interactive Buildings

DGX agent

arXiv:2605.28168v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated promising capability in generating reward functions for deep reinforcement learning (DRL)-based building

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models

DGX agent

arXiv:2503.01829v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) demonstrate persuasive capabilities that rival human-level persuasion. While these capabilities can be used for s

model-releasesarxiv-cs-ai
28 May 2026
Safety

Completely agree, @hwchase17 ! This is the meta-level breakthrough we've all been waiting for. Self-optimizing loops finally feel production…

DGX agent

Completely agree, @hwchase17 ! This is the meta-level breakthrough we've all been waiting for. Self-optimizing loops finally feel production-ready because LangSmith Engine turns evaluation from a manu

safetyharrison-chase--x
27 May 2026
Tutorials

I used the N.E.A.T algorithm to teach AI how to control a worm in my game in making! It uses evolution to improve. [P]

DGX agent

The N.E.A.T (NeuroEvolution of Augmenting Topologies) algorithm is an evolutionary machine learning approach that evolves neural networks to solve control problems. This post describes applying N.E.A.

tutorialsr-machinelearning
27 May 2026
Safety

Intelligent Offloading in Vehicular Edge Computing: A Comprehensive Review of Deep Reinforcement Learning Approaches and Architectures

DGX agent

arXiv:2502.06963v3 Announce Type: replace-cross Abstract: The increasing complexity of Intelligent Transportation Systems (ITS) has led to significant interest in computational offloading to external

safetyarxiv-cs-ai
27 May 2026
Model Releases

Traceable Knowledge Graph Reasoning Enables LLM-Assisted Decision Support for Industrial VOCs in the Steel Industry

DGX agent

arXiv:2605.27071v1 Announce Type: new Abstract: Key knowledge for steel-industry volatile organic compounds (VOCs) governance is scattered across unstructured scientific literature, making it difficul

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

When Does Deep RL Beat Calibrated Baselines? A Benchmark Study on Adaptive Resource Control

DGX agent

arXiv:2605.26418v1 Announce Type: cross Abstract: A properly calibrated rule-based autoscaler can beat every one of six mainstream deep reinforcement learning (DRL) algorithms on cost across every wor

model-releasesarxiv-cs-ai
27 May 2026
← Previous
1…309310311312313…371
Next →