AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “research”

GridTimelineEvolution
25,629 results
20 May 2026

Precision Physical Activity Prescription via Reinforcement Learning for Functional Actions

SafetyDGX agent

arXiv:2605.19208v1 Announce Type: cross Abstract: Physical activity (PA) plays an important role in maintaining and improving health. Daily steps have been a key PA measure that is easily accessible w

PromptRad: Knowledge-Enhanced Multi-Label Prompt-Tuning for Low-Resource Radiology Report Labeling

Model ReleasesDGX agent

arXiv:2605.20052v1 Announce Type: cross Abstract: Automatic report labeling facilitates the identification of clinical findings from unstructured text and enables large-scale annotation for medical im

Quantifying the Generalization Gap in Seizure Detection: A Large-Scale Empirical Benchmark via the SzCORE Challenge

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2505.18191v2 Announce Type: replace-cross Abstract: Reliable automatic seizure detection from long-term electroencephalography (EEG) remains an unsolved challenge, as current models often fail t

Quoting SpaceX S-1

ToolsDGX agent

We have the ability to use compute resources to support our proprietary AI applications (such as Grok 5, which is currently being trained at COLOSSUS II), while also providing access to select compute

Rapid patient-specific neural networks for intraoperative X-ray to volume registration

SafetyDGX agent

arXiv:2503.16309v2 Announce Type: replace-cross Abstract: Advanced navigation techniques in image-guided interventions and surgical robotics require the rapid and precise alignment of 3D preoperative

Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior

Model ReleasesDGX agent

arXiv:2510.14261v2 Announce Type: replace Abstract: We present an experimental recipe for studying the relationship between training data and language model (LM) behavior. We outline steps for interve

Robotics-Inspired Guardrails for Foundation Models in Socially Sensitive Domains

SafetyDGX agent

arXiv:2605.19940v1 Announce Type: new Abstract: Foundation models are increasingly deployed in socially sensitive domains such as education, mental health, and caregiving, where failures are often cum

Self-improving AI is a big deal! As a first step, I've been exploring how much of the post-training can be automated. Here is a first post o…

Model ReleasesDGX agent

Self-improving AI is a big deal! As a first step, I've been exploring how much of the post-training can be automated. Here is a first post on how I am using @FireworksAI_HQ Agent to automate LLM fine-

Smartphone-based Circular Plot Sampling for Forest Inventory

TutorialsDGX agent

arXiv:2605.19213v1 Announce Type: new Abstract: Circular sample plots are a cornerstone of forest inventory, yet accurate measurement of tree diameter at breast height (DBH) and spatial location withi

Sonar-TS: Search-Then-Verify Natural Language Querying for Time Series Databases

Model ReleasesDGX agent

arXiv:2602.17001v2 Announce Type: replace Abstract: Natural Language Querying for Time Series Databases (NLQ4TSDB) aims to assist non-expert users retrieve meaningful events, intervals, and summaries

Stoked to work closely with @Vtrivedy10 on this. I've spoken about it before but an emerging trend i'm seeing with many of the ai-native com…

AgentsDGX agent

Stoked to work closely with @Vtrivedy10 on this. I've spoken about it before but an emerging trend i'm seeing with many of the ai-native companies I work with (shoutout @larsen_weigle_ ) is a focus on

Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits

AgentsDGX agent

arXiv:2605.18890v1 Announce Type: cross Abstract: The scientific claims drawn from LLM social simulations should be no stronger than the robustness audits that support them. Generative agents bring ne

Synthetic Data Generation for Brain-Computer Interfaces: Overview, Benchmarking, and Future Directions

Model ReleasesDGX agent

arXiv:2603.12296v2 Announce Type: replace-cross Abstract: Deep learning has achieved transformative performance across diverse domains, largely driven by large-scale and high-quality training data. In

The Accessibility Capability Boundary: Operational Limits and Expansion Potential of AI-Generated Browser-Native Accessibility Systems

SafetyDGX agent

arXiv:2605.19638v1 Announce Type: cross Abstract: As large language models (LLMs) demonstrate increasing competence in synthesizing functional user interfaces, a fundamental question emerges in access

Toward an AI-Powered Computational Testbed for Workforce Policy

SafetyDGX agent

arXiv:2605.19064v1 Announce Type: cross Abstract: Workforce transformations are difficult to forecast and costly to mismanage. In particular, the integration of artificial intelligence into knowledge

Trajectory Planning and Control near the Limits: an Open Experimental Benchmark on the RoboRacer Platform

Model ReleasesDGX agent

arXiv:2605.19881v1 Announce Type: new Abstract: We present a modular framework to benchmark new and existing methods for trajectory planning and control in high-acceleration maneuvers that push autono

When Web Apps Heal Themselves: A MAPE-K Based Approach to Fault Tolerance and Adaptive Recovery

AgentsDGX agent

arXiv:2605.19261v1 Announce Type: cross Abstract: Ensuring the reliability and resilience of modern web applications remains a critical challenge due to increasing system complexity and dynamic runtim

ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Model ReleasesDGX agent

arXiv:2505.04588v3 Announce Type: replace Abstract: Effective information searching is essential for enhancing the reasoning and generation capabilities of large language models (LLMs). Recent researc

19 May 2026

4DLidarOpen: An Open 4D FMCW Lidar Dataset for Motion-Aware Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.18074v1 Announce Type: new Abstract: We present 4DLidarOpen, a large-scale open multi-modal dataset for autonomous driving, centered on 4D frequency-modulated continuous-wave (FMCW) Lidar s

A Feature-Driven Framework for Software Fault Prediction

Model ReleasesDGX agent

arXiv:2605.17611v1 Announce Type: cross Abstract: Software fault prediction (SFP) is a critical task in software engineering, enabling early identification of faults in modules to improve software qua

A Pilot Benchmark for NL-to-FOL Translation in Planetary Exploration

Model ReleasesDGX agent

arXiv:2605.17911v1 Announce Type: new Abstract: Future planetary exploration envisions autonomous robotic agents operating under severe communication constraints, without global positioning, and with

Active Budget Allocation for Efficient Scaling Law Estimation via Surrogate-Guided Pruning

ApplicationsDGX agent

arXiv:2605.17234v1 Announce Type: new Abstract: Predicting model performance at larger scales enables the design of training strategies and architectures tailored to specific performance targets. Empi

AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent

AgentsDGX agent

arXiv:2602.03955v2 Announce Type: replace Abstract: While large language model (LLM) multi-agent systems achieve superior reasoning performance through iterative debate, practical deployment is limite

AI Agents May Always Fall for Prompt Injections

SafetyDGX agent

arXiv:2605.17634v1 Announce Type: cross Abstract: Prompt injection is the most critical vulnerability in deployed AI agents. Despite recent progress, we show that the prevailing defense paradigm (data

Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework

SafetyDGX agent

arXiv:2605.16516v1 Announce Type: cross Abstract: Long-term interaction with LLM-based systems may produce alignment drift: a gradual process in which system outputs become less constrained by the use

ARROW: Augmented Replay for RObust World models

SafetyDGX agent

arXiv:2603.11395v2 Announce Type: replace-cross Abstract: Continual reinforcement learning challenges agents to acquire new skills while retaining previously learned ones with the goal of improving pe

Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling

Model ReleasesDGX agent

arXiv:2605.17971v1 Announce Type: cross Abstract: Despite rigorous safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. Existing black-box methods often rely on heuri

Bench2Drive-Robust: Benchmarking Closed-Loop Autonomous Driving under Deployment Perturbations

Model ReleasesDGX agent

arXiv:2605.18059v1 Announce Type: new Abstract: Robustness is a critical requirement for deploying autonomous driving systems in the real world. Existing robustness benchmarks for autonomous driving h

Beyond Compliance: How AI Could Help Creative Writers by Refusing Them

SafetyDGX agent

arXiv:2605.16272v1 Announce Type: cross Abstract: Mainstream creativity support design prioritizes compliant AI for seamless writing interactions, but concerns over inappropriate AI reliance highlight

Breaking Annotation Barriers: Generalized Video Quality Assessment via Ranking-based Self-Supervision

Model ReleasesDGX agent

arXiv:2505.03631v4 Announce Type: replace Abstract: Video quality assessment (VQA) is essential for quantifying perceptual quality in various video processing workflows, spanning from camera capture s

Can LLMs Generate and Solve Linguistic Olympiad Puzzles?

Model ReleasesDGX agent

arXiv:2509.21820v2 Announce Type: replace Abstract: In this paper, we introduce a combination of novel and exciting tasks: the solution and generation of linguistic puzzles. We focus on puzzles used i

CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning

Model ReleasesDGX agent

arXiv:2605.17176v1 Announce Type: new Abstract: Emotion understanding is a core capability for LLMs to interact effectively with humans, yet existing evaluation paradigms rely on discrete emotion labe

CayleyPy RL: Pathfinding and Reinforcement Learning on Cayley Graphs

Model ReleasesDGX agent

arXiv:2502.18663v3 Announce Type: replace Abstract: This paper is the second in a series of studies on developing efficient artificial intelligence-based approaches to pathfinding on extremely large g

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting

Model ReleasesDGX agent

arXiv:2508.04227v2 Announce Type: replace Abstract: Vision-language models (VLMs) and the recent surge of Multimodal Large Language Models (MLLMs) have revolutionized artificial intelligence with unpr

CooT: Learning to Coordinate In-Context with Coordination Transformers

Model ReleasesDGX agent

arXiv:2506.23549v3 Announce Type: replace Abstract: Effective coordination among unfamiliar partners remains a major challenge in multi-agent systems. Existing approaches, such as population-based met

Deep sequence models tend to memorize geometrically; it is unclear why

SafetyDGX agent

arXiv:2510.26745v3 Announce Type: replace-cross Abstract: Deep sequence models are said to store atomic facts predominantly in the form of associative memory: a brute-force lookup of co-occurring enti

EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning

Model ReleasesDGX agent

arXiv:2605.17262v1 Announce Type: new Abstract: Despite extensive efforts on egocentric video datasets and benchmarks, understanding users' internal states, which is crucial for enabling seamless AI a

Estimating Item Difficulty with Large Language Models as Experts

SafetyDGX agent

arXiv:2605.18562v1 Announce Type: cross Abstract: Accurate estimates of item difficulty are essential for valid assessment and effective adaptive learning. However, for newly created tasks, response d

Evaluating Cognitive Age Alignment in Interactive AI Agents

Model ReleasesDGX agent

arXiv:2605.17894v1 Announce Type: new Abstract: While agentic AI and its core multimodal large language models (MLLMs) have demonstrated remarkable promise in language and visual reasoning across doma

Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs

Model ReleasesDGX agent

arXiv:2511.12710v2 Announce Type: replace Abstract: Automated red teaming frameworks for Large Language Models (LLMs) have become increasingly sophisticated, yet many still formulate attack optimizati

EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective

Model ReleasesDGX agent

arXiv:2605.18421v1 Announce Type: cross Abstract: Recent benchmarks for Large Language Model (LLM) agents mainly evaluate reasoning, planning, and execution. However, memory is also essential for agen

From Demographics to Survey Anchors: Evaluating LLM Agents for Modeling Retirement Attitudes

SafetyDGX agent

arXiv:2605.16303v1 Announce Type: cross Abstract: Large language models (LLM) agents may offer tools to predict human responses to surveys. A common technique for defining these agents uses only demog

Generative AI Advertising as a Problem of Trustworthy Commercial Intervention

AgentsDGX agent

arXiv:2605.18673v1 Announce Type: cross Abstract: Major deployed generative AI advertising systems preserve a visible boundary between commercial content and AI-generated responses. Yet empirical rese

Generative Artificial Intelligence for Literature Reviews

Model ReleasesDGX agent

arXiv:2605.16475v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI), based on large-language models (LLMs), such as ChatGPT, has taken organizations, academia, and the public

Google just redesigned the search box for the first time in 25 years — here’s why it matters more than you think.

Model ReleasesDGX agent

For a quarter century, the Google search box has been one of the most recognizable interfaces in computing: a thin white rectangle, a blinking cursor, a few typed words, and a list of blue links. On T

Graph Embedding in the Graph Fractional Fourier Transform Domain

Model ReleasesDGX agent

arXiv:2508.02383v2 Announce Type: replace Abstract: Spectral graph embedding plays a critical role in graph representation learning by generating low-dimensional vector representations from graph spec

GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games

Model ReleasesDGX agent

arXiv:2508.08501v3 Announce Type: replace Abstract: We introduce GVGAI-LLM, a video game benchmark for evaluating the reasoning and problem-solving capabilities of large language models (LLMs). Built

HalluScore: Large Language Model Hallucination Question Answering Benchmark

Model ReleasesDGX agent

arXiv:2605.17007v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable progress in natural language generation, but remain susceptible to hallucination. In response to g

High-Resolution Reference Image Assisted Volumetric Super-Resolution of Cardiac Diffusion Weighted Imaging

TutorialsDGX agent

arXiv:2310.20389v2 Announce Type: replace-cross Abstract: Diffusion Tensor Cardiac Magnetic Resonance (DT-CMR) is the only in vivo method to non-invasively examine the microstructure of the human hear

HTSC-2025: A Benchmark Dataset of Ambient-Pressure High-Temperature Superconductors for AI-Driven Critical Temperature Prediction

Model ReleasesDGX agent

arXiv:2506.03837v2 Announce Type: replace-cross Abstract: The discovery of high-temperature superconducting materials holds great significance for human industry and daily life. In recent years, resea

I built a tool that shows you what GPT-2 is 'thinking' in real-time as it generates 3D graph of concept activations per token

IndustryDGX agent

This post describes a visualization tool that provides real-time insight into GPT-2's internal processing by displaying concept activations as a 3D graph, with each token generation showing which neur

Introducing OpenAI for Singapore

Model ReleasesDGX agent

OpenAI announced its expansion into Singapore, establishing a presence in the Asia-Pacific region to better serve local customers and developers. The announcement likely covers OpenAI's commitment to

Inventorship in AI-Assisted Inventions: Designing an Experiment to Shape Case Law

TutorialsDGX agent

arXiv:2605.16528v1 Announce Type: cross Abstract: The latest improvements in artificial intelligence (AI) raise new challenges for intellectual property laws, particularly concerning the inventorship

iPOE: Interpretable Prompt Optimization via Explanations

TutorialsDGX agent

arXiv:2605.18113v1 Announce Type: new Abstract: Prompt optimization has often been framed as a discrete search problem to find high-performing and robust instructions for an LLM. However, the search r

KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy

Model ReleasesDGX agent

arXiv:2605.16439v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have emerged as a critical and fast-growing extension of Large Language Models (LLMs) that enable multimodal reasoning t

Latency-Aware Deep Learning Benchmark for Real-Time Cyber-Physical Attack and Fault Classification in Inverter-Dominated Power Grids

Model ReleasesDGX agent

arXiv:2605.17256v1 Announce Type: cross Abstract: This work introduces a latency-aware benchmarking framework for evaluating deep learning models in power system anomaly detection using high-fidelity,

Learning What Evaluators Value: A Reliable Approach to Modeling Evaluator Preferences

TutorialsDGX agent

arXiv:2605.16615v1 Announce Type: new Abstract: In many applications, human and LLM evaluators use assessments of relevant criteria to create an overall evaluation for an item or individual. For examp

Linguistic Uncertainty and Reply Engagement on X: A Cross-Domain Replication of the Uncertainty-Reply Asymmetry

SafetyDGX agent

arXiv:2605.16289v1 Announce Type: cross Abstract: Linguistic uncertainty is common in social media, but its relationship with engagement remains unclear across languages and topics. Using 2,258 Englis

LiTS: A Modular Framework for LLM Tree Search

Model ReleasesDGX agent

arXiv:2603.00631v2 Announce Type: replace Abstract: LiTS is a modular Python framework for LLM reasoning via tree search. It decomposes tree search into three reusable components (Policy, Transition,

Low Latency Gaze Tracking via Latent Optical Sensing

ApplicationsDGX agent

arXiv:2605.17990v1 Announce Type: new Abstract: We present a real-time gaze tracking system that directly acquires task-relevant latent features using a fully passive optical encoder. Instead of formi

← Previous
1…402403404405406…428
Next →