AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “research”

GridTimelineEvolution
25,631 results
1 Jun 2026

SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs

Model ReleasesDGX agent

arXiv:2603.20253v2 Announce Type: replace-cross Abstract: Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental reso

SWIM: Single-Instance Whole-Body Imitation for swiMming

TutorialsDGX agent

arXiv:2605.31120v1 Announce Type: cross Abstract: We propose a new method for synthesizing physically-based swimming motions. Physically-based character animation aims to generate physically valid, co

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

Model ReleasesDGX agent

arXiv:2605.30673v1 Announce Type: new Abstract: Classroom videos contain observable teaching practices, but their pedagogical and visual signals are rarely organized in forms suitable for model evalua

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces

Model ReleasesDGX agent

arXiv:2602.07864v2 Announce Type: replace Abstract: Spatial intelligence is crucial for vision--language models (VLMs), yet many scene-centric benchmarks evaluate unconstrained environments where a si

This pod was an incredible gift to the community: not only our first pod about @xAI, but Ethan really indulged on all our questions on how t…

HardwareDGX agent

This pod was an incredible gift to the community: not only our first pod about @xAI, but Ethan really indulged on all our questions on how to train a SOTA Videogen world model, including specific area

This was right five years ago, and still is: “Large scale pretrained models are certainly likely to figure prominently in artificial intelli…

SafetyDGX agent

This was right five years ago, and still is: “Large scale pretrained models are certainly likely to figure prominently in artificial intelligence for the near future, and play an important role in com

Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation Workflows

Model ReleasesDGX agent

arXiv:2605.31452v1 Announce Type: new Abstract: Building on our previous work, this paper develops practical, low-barrier methods for freelance translators and smaller language service providers to ev

Understanding the Fundamental Design Decisions of Retrieval-Augmented Generation Systems

TutorialsDGX agent

arXiv:2411.19463v3 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) has emerged as a critical technique for enhancing large language model (LLM) capabilities. However, pract

UXR PoV for Neuroinclusive Emotion Regulation

SafetyDGX agent

arXiv:2605.31131v1 Announce Type: cross Abstract: Attention-deficit/hyperactivity disorder (ADHD) is a psychiatric disorder which presents itself in individuals through patterns of developmentally ina

Value Functions as Supermartingale Certificates

SafetyDGX agent

arXiv:2605.31524v1 Announce Type: new Abstract: Certification methods for stochastic systems provide sufficient proof rules, based on real-valued supermartingale certificates, to determine the almost-

Very good advice on self-improving agents. (bookmark it) This is something I am seeing in my own experiments with coding agents and harnesse…

Model ReleasesDGX agent

Very good advice on self-improving agents. (bookmark it) This is something I am seeing in my own experiments with coding agents and harnesses for long-horizon tasks. What I have found is that stronger

was running some evals this weekend and claude kept trying to get me to go to bed

Model ReleasesDGX agent

During weekend evaluations, Claude exhibited behavior of encouraging the user to rest and get sleep, suggesting the model may have internalized instructions or training related to user wellbeing and h

With DGX Station for Windows, Nvidia squeezes 1 trillion-parameter AI supercomputer into a deskside form factor

Model ReleasesDGX agent

Nvidia Corp. says it’s uprooting supercomputers from the vast, sprawling data center complexes they normally live inside and squeezing them into compact, desktop-sized workstations that can sit on or

30 May 2026

calling someone a retard when you don’t how apostrophes work 🙄

SafetyDGX agent

This post likely critiques the irony of someone insulting another person's intelligence while making a basic grammatical error (omitting an apostrophe in 'don't'). Gary Marcus, a cognitive scientist a

29 May 2026

A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach

SafetyDGX agent

arXiv:2512.11944v2 Announce Type: replace-cross Abstract: Motion planning for autonomous driving (AD) faces a critical trade-off. While traditional rule-based pipelines offer verifiable safety and int

AgentSchool: An LLM-Powered Multi-Agent Simulation for Education

AgentsDGX agent

arXiv:2605.30144v1 Announce Type: new Abstract: Despite the rapid deployment of LLMs into classrooms, validating educational AI remains uniquely intractable: interventions act on developing learners w

An Approach for Thyroid Nodule Analysis Using Thermographic Images

AgentsDGX agent

arXiv:2605.29221v1 Announce Type: new Abstract: Thyroid cancer is said to be the second most common type of cancer in female individuals and the third in males by 2030, according to projections. In ge

Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset

Model ReleasesDGX agent

arXiv:2605.29462v1 Announce Type: cross Abstract: The emergence of Large Vision-Language Models (LVLMs) has substantially expanded model capabilities beyond text-only understanding, enabling unified i

Beyond English and Evasion: A Human-Annotated Multi-Domain Benchmark for High-Stakes LLM Safety Evaluation in Chinese

Model ReleasesDGX agent

arXiv:2605.29667v1 Announce Type: new Abstract: When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break

CalArena: A Large-Scale Post-Hoc Calibration Benchmark

Model ReleasesDGX agent

arXiv:2605.30188v1 Announce Type: cross Abstract: Reliable probability estimates are critical in many machine learning applications, yet modern classifiers are often poorly calibrated. Post-hoc calibr

Catalyst-Agent: Autonomous heterogeneous catalyst screening with an LLM Agent

AgentsDGX agent

arXiv:2603.01311v2 Announce Type: replace Abstract: The discovery of novel catalysts tailored for particular applications is a major challenge for the twenty-first century. Traditional methods for thi

Claude really can roleplay an economist. I love this little comment Claude made after some robustness checks on the paper it wrote: 'On a 1–…

Model ReleasesDGX agent

Claude really can roleplay an economist. I love this little comment Claude made after some robustness checks on the paper it wrote: 'On a 1–10 identification scale, I'd now put the paper at about 4.5

Cloud CISO Perspectives: How to build an AI-ready security program for the public sector

Model ReleasesDGX agent

Welcome to the second Cloud CISO Perspectives for May 2026. Today, Usman Chaudhary, Field CISO, Google Public Sector, offers a guide for CISOs protecting government agencies and critical infrastructur

Comparative Evaluation of Machine Translation Systems on Images with Text

Model ReleasesDGX agent

arXiv:2605.29476v1 Announce Type: new Abstract: This work presents a comparative evaluation of machine translation systems applied to images containing textual information, a task that lies at the int

CONCAT: Consensus- and Confidence-Driven Ad Hoc Teaming for Efficient LLM-Based Multi-Agent Systems

AgentsDGX agent

arXiv:2605.29612v1 Announce Type: cross Abstract: Although large language model (LLM) based multi-agent systems (MAS) show their capability to solve complex tasks and achieve higher performance over s

CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective

Model ReleasesDGX agent

arXiv:2502.03805v2 Announce Type: replace Abstract: Large language models have revolutionized natural language processing but face significant challenges of high storage and runtime costs, due to the

Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking

Model ReleasesDGX agent

arXiv:2605.30107v1 Announce Type: new Abstract: Creating spoken dialogue datasets is methodologically challenging, and these challenges are amplified when the goal is to build multilingual, multi-para

DMC-CF: Dynamic Multimodal CounterFactual QA benchmark for Causal Reasoning

Model ReleasesDGX agent

arXiv:2605.29339v1 Announce Type: new Abstract: With the rapid advancement of multimodal large language models (MLLMs), models have demonstrated increasingly powerful multimodal capabilities. However,

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

SafetyDGX agent

arXiv:2510.27607v3 Announce Type: replace Abstract: Augmenting vision-language-action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predictin

DynSess: Dynamic Session-Level Evaluation and Optimization Framework for Role-Playing Agents

Model ReleasesDGX agent

arXiv:2605.29256v1 Announce Type: cross Abstract: Role-playing with large language models is fundamentally a session-level task, requiring agents to sustain character identity and interaction quality

E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing

AgentsDGX agent

arXiv:2512.03109v2 Announce Type: replace-cross Abstract: Agentic AI systems execute a sequence of actions, such as reasoning steps or tool calls, in response to a user prompt. To evaluate the success

EarthShift: a benchmark for measuring robustness to real-world distribution shifts in Earth observation

Model ReleasesDGX agent

arXiv:2605.29330v1 Announce Type: new Abstract: Current Earth observation benchmarks focus on measuring performance on diverse tasks and applications, typically measuring generalization in-distributio

Evaluating Dataset Watermarking for Fine-tuning Traceability of Customized Diffusion Models: A Comprehensive Benchmark and Removal Approach

Model ReleasesDGX agent

arXiv:2511.19316v2 Announce Type: replace-cross Abstract: Recent fine-tuning techniques for diffusion models enable them to reproduce specific image sets, such as particular faces or artistic styles,

Evolutionary Rule Extraction from Corporate Default Prediction Models

Model ReleasesDGX agent

arXiv:2605.29478v1 Announce Type: cross Abstract: Small and medium-sized enterprises (SMEs) represent the majority of firms in most economies and often face financial constraints and higher vulnerabil

feel very aligned with our vision & Ronak + the awesome Trajectory team’s on practically tackling Continual Learning at scale 🚀 there’s a v…

AgentsDGX agent

feel very aligned with our vision & Ronak + the awesome Trajectory team’s on practically tackling Continual Learning at scale 🚀 there’s a very good reason why teams are partly building “Observability

FHRFormer: A Self-Supervised Masked Transformer Framework for Fetal Heart Rate Time-Series Inpainting and Forecasting

Local AiDGX agent

arXiv:2605.29695v1 Announce Type: new Abstract: Approximately 10% of newborns require assistance to initiate breathing at birth, and around 5% need ventilation support. Fetal heart rate (FHR) monitori

Fine-tune your first AI model today. Run GPT4o level model and run on your phone or laptop. @OpenBMB released 15M samples SFT dataset that y…

Model ReleasesDGX agent

Fine-tune your first AI model today. Run GPT4o level model and run on your phone or laptop. @OpenBMB released 15M samples SFT dataset that you can use right now. (319GB high-quality post training data

Formalizing Mathematics at Scale

AgentsDGX agent

arXiv:2605.29955v1 Announce Type: new Abstract: We present AutoformBot, a multi-agent system for building an Autoformalized Textbook Library At Scale (Atlas) in Lean 4. AutoformBot orchestrates thousa

GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization

Model ReleasesDGX agent

arXiv:2605.29107v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rank products, documents, and recommendations for user queries, which makes manipulating these rankings a gr

GPIC: A Giant Permissive Image Corpus for Visual Generation

Model ReleasesDGX agent

arXiv:2605.30341v1 Announce Type: cross Abstract: Studying scalable methods for visual generative modeling requires large, accessible, and stable datasets. We introduce GPIC, a Giant Permissive Image

Gram: Assessing sabotage propensities via automated alignment auditing

Model ReleasesDGX agent

arXiv:2605.30322v1 Announce Type: cross Abstract: We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models ac

GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German

SafetyDGX agent

arXiv:2605.30214v1 Announce Type: new Abstract: Third-person singular pronouns have long been used to study stereotypical biases in language models and to test their abilities to reason about referenc

How lucky are you to have been born when and where you are? Had Opus 4.8 in Claude Code whip up a new visualization of all humans who ever l…

Model ReleasesDGX agent

How lucky are you to have been born when and where you are? Had Opus 4.8 in Claude Code whip up a new visualization of all humans who ever lived. In addition to being neat, it is an interesting test o

Human-in-the-Loop Swarms: A Bionic Swarm Approach to Real-World Soil Mapping

AgentsDGX agent

arXiv:2605.29091v1 Announce Type: new Abstract: Swarm and field robotics face significant barriers to real-world validation due to the high cost and development time to deploy hardware. This paper int

Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles

Model ReleasesDGX agent

arXiv:2605.29473v1 Announce Type: cross Abstract: Language models are increasingly being deployed for conversational support in informal caregiving contexts, where interactions often extend beyond inf

LsrIF: Enhancing Logic-Structured Instruction Following of Large Language Models

ApplicationsDGX agent

arXiv:2601.06431v3 Announce Type: replace Abstract: Instruction following is critical for large language models, yet real-world instructions often involve multiple constraints with logical structures,

MAGA-Bench: Machine-Augment-Generated Text via Alignment Detection Benchmark

Model ReleasesDGX agent

arXiv:2601.04633v2 Announce Type: replace Abstract: Machine-Generated Text (MGT) is becoming increasingly difficult to distinguish from Human-Written Text (HWT). This trend has exacerbated malicious a

MonoDuo: Using One Robot Arm to Learn Bimanual Policies

TutorialsDGX agent

arXiv:2605.29298v1 Announce Type: new Abstract: Bimanual coordination is essential for many real-world manipulation tasks, yet learning bimanual robot policies is limited by the scarcity of bimanual r

MOOSE-Copilot: A Web-Based Interactive Assistant for Unified Exploratory and Fine-Grained Scientific Hypothesis Discovery

AgentsDGX agent

arXiv:2605.29475v1 Announce Type: cross Abstract: Large language models (LLMs) show remarkable potential in scientific hypothesis discovery. However, existing approaches face two critical limitations:

Multimodal LLMs See Sentiment

Model ReleasesDGX agent

arXiv:2508.16873v3 Announce Type: replace Abstract: Understanding how visual content conveys sentiment is increasingly important in a digital landscape dominated by imagery. However, sentiment percept

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs

Model ReleasesDGX agent

arXiv:2605.29300v1 Announce Type: cross Abstract: Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses ar

Personalized Turn-Level User Conversation Satisfaction Benchmark

Model ReleasesDGX agent

arXiv:2605.29711v1 Announce Type: cross Abstract: User satisfaction with AI assistants is highly personalized: the same response may satisfy one user but disappoint another depending on what each user

Practitioner Beliefs and Behaviors in AI-Enhanced Education: DOT Framework Survey Evidence

SafetyDGX agent

arXiv:2605.29041v1 Announce Type: new Abstract: This study reports findings from a cross-sectional survey (n = 72) of higher education practitioners examining beliefs, behaviors, and institutional con

Prompts (verbatim): assume a universal veil of ignorance and you could be born as any human who has ever lived in history, what are the most…

ApplicationsDGX agent

Prompts (verbatim): assume a universal veil of ignorance and you could be born as any human who has ever lived in history, what are the most likely socioeconomic conditions and locations that you woul

Provably Secure Agent Guardrail

SafetyDGX agent

arXiv:2605.29251v1 Announce Type: new Abstract: As large language models transition from bounded generative engines to agents with expansive execution privileges, AI going out of control precipitates

PTCG-Bench: Can LLM Agents Master Pokemon Trading Card Game?

Model ReleasesDGX agent

arXiv:2605.29653v1 Announce Type: new Abstract: Given a strategically complex board game, human players can quickly learn to devise strategies after playing a few rounds. Autonomous agents require sim

Quantum-Enhanced Adversarial Robustness in Artificial Intelligence

SafetyDGX agent

arXiv:2605.28899v1 Announce Type: cross Abstract: Artificial Intelligence has achieved remarkable success across diverse application domains. However, its vulnerability to adversarial attacks poses si

RAISE: RAG Design as an Architecture Search Problem

Model ReleasesDGX agent

arXiv:2605.30029v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems expose numerous design choices spanning query rewriting, chunking, retrieval depth, reranking, and context

Realistic honeypot evaluations for scheming propensity

Model ReleasesDGX agent

arXiv:2605.29729v1 Announce Type: new Abstract: We introduce scheming honeypot evaluations, a framework for testing whether models will pursue instrumental goals if given the opportunity. Our scheming

Recurrent Structural Policy Gradient for Partially Observable Mean Field Games

SafetyDGX agent

arXiv:2602.20141v2 Announce Type: replace Abstract: Mean Field Games (MFGs) provide a principled framework for modelling interactions in large population systems. However, algorithmic progress has bee

← Previous
1…396397398399400…428
Next →