AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “automated”

GridTimelineEvolution
4,952 results
29 Apr 2026

Exploring Remote Photoplethysmography for Neonatal Pain Detection from Facial Videos

Model ReleasesDGX agent

arXiv:2604.25680v1 Announce Type: new Abstract: Unaddressed pain in neonates can lead to adverse effects, including delayed development and slower weight gain, emphasising the need for more objective

Faithfulness-QA: A Counterfactual Entity Substitution Dataset for Training Context-Faithful RAG Models

Model ReleasesDGX agent

arXiv:2604.25313v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) models frequently produce answers grounded in parametric memory rather than the retrieved context, undermining the

GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.24929v1 Announce Type: new Abstract: Agent benchmarks remain largely English-centric, while their multilingual versions are often built with machine translation (MT) and limited post-editin

GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment

Model ReleasesDGX agent

arXiv:2604.25370v1 Announce Type: new Abstract: The release of GPT-image-2 by OpenAI marks a watershed moment in AI-generated imagery: the boundary between photographic reality and synthetic content h

GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning

ResearchDGX agent

arXiv:2604.25459v1 Announce Type: new Abstract: Embodied AI research is undergoing a shift toward vision-centric perceptual paradigms. While massively parallel simulators have catalyzed breakthroughs

Heterogeneous Variational Inference for Markov Degradation Hazard Models: Discretized Mixture with Interpretable Clusters

ApplicationsDGX agent

arXiv:2604.24818v1 Announce Type: new Abstract: Bayesian finite mixture models can identify discrete risk clusters (low-risk vs. high-risk equipment), but face three critical bottlenecks: (1) insuffic

Inside Elevance Health’s push to keep humans at the center of AI-driven care

ApplicationsDGX agent

As AI transforms healthcare operations, the industry’s most consequential challenge has shifted from building models to deploying them responsibly in regulated workflows where trust, accuracy and huma

Measuring the Sensitivity of Classification Models with the Error Sensitivity Profile

ResearchDGX agent

arXiv:2604.25765v1 Announce Type: new Abstract: The quality of training data is critical to the performance of machine learning models. In this paper, the Error Sensitivity Profile (ESP) is proposed.

New York-based Actively, which offers AI sales agents for account management, raised a 45M Series B co-led by TCV and First Harmonic at a 250M valuation (Sofia Chierchio/Forbes)

IndustryDGX agent

Sofia Chierchio / Forbes: New York-based Actively, which offers AI sales agents for account management, raised a 45M Series B co-led by TCV and First Harmonic at a 250M valuation — Actively AI has rai

Use of What-if Scenarios to Help Explain Artificial Intelligence Models for Neonatal Health

ResearchDGX agent

arXiv:2410.09635v2 Announce Type: replace Abstract: Early detection of intrapartum risks enables timely interventions to prevent or mitigate adverse labor outcomes such as cerebral palsy. However, acc

28 Apr 2026

A Digital Pathology Resource for Liver Cancer Quantification with Datasets, Benchmarks, and Tools

Local AiDGX agent

arXiv:2604.22858v1 Announce Type: new Abstract: Liver cancer, especially hepatocellular carcinoma (HCC), imposes a substantial global disease burden. Accurate diagnosis and prognostic assessment direc

A Multi-Dimensional Audit of Politically Aligned Large Language Models

SafetyDGX agent

arXiv:2604.24429v1 Announce Type: new Abstract: As the application of Large Language Models (LLMs) spreads across various industries, there are increasing concerns about the potential for their misuse

Accelerating New Product Introduction for Visual Quality Inspection via Few-Shot Diffusion-Based Defect Synthesis

ResearchDGX agent

arXiv:2604.22850v1 Announce Type: new Abstract: Industrial visual inspection systems often suffer from a severe scarcity of labeled defect data, particularly during the early stages of New Product Int

AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking

AgentsDGX agent

arXiv:2604.23581v1 Announce Type: cross Abstract: Agentic systems that chain reasoning, tool use, and synthesis into multi-step workflows are entering production, yet prevailing evaluation practices l

Agentic Witnessing: Pragmatic and Scalable TEE-Enabled Privacy-Preserving Auditing

Model ReleasesDGX agent

arXiv:2604.24203v1 Announce Type: cross Abstract: Auditing the semantic properties of proprietary data creates a fundamental tension: verification requires transparent access, while proprietary rights

Algorithmic Administration and the EU AI Act: Legal Principles for Public Sector Use of AI

SafetyDGX agent

arXiv:2604.22765v1 Announce Type: cross Abstract: The increasing use of artificial intelligence (AI) by public authorities introduces both opportunities for innovation and significant challenges for t

An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code

Model ReleasesDGX agent

arXiv:2604.23361v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong performance on a wide range of software engineering tasks, including code generation and analysi

AnemiaVision: Non-Invasive Anemia Detection via Smartphone Imagery Using EfficientNet-B3 with TrivialAugmentWide, Mixup Augmentation, and Persistent Patient History Management

SafetyDGX agent

arXiv:2604.22964v1 Announce Type: new Abstract: Anemia affects over one billion people globally and remains severely under-diagnosed in low-resource regions where laboratory blood tests are inaccessib

b8966

Local AiDGX agent

b8966 is one of the latest releases in the llama.cpp project , a C/C++ implementation for efficient LLM inference. This release represents an intermediate development build in the llama.cpp repository

Built In, Not Bolted On: What AI-Native Actually Means in Cybersecurity

IndustryDGX agent

AI-native cybersecurity refers to security systems designed with artificial intelligence as a fundamental architectural component rather than as an added afterthought, enabling more effective threat d

Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters

AgentsDGX agent

arXiv:2604.24710v1 Announce Type: new Abstract: Objective. Clinical AI documentation systems require evaluation methodologies that are clinically valid, economically viable, and sensitive to iterative

Cataract-LMM Large-Scale Multi-Source Multi-Task Benchmark for Deep Learning in Surgical Video Analysis

Model ReleasesDGX agent

arXiv:2510.16371v2 Announce Type: replace-cross Abstract: The development of computer-assisted surgery systems relies on large-scale, annotated datasets. Existing cataract surgery resources lack the d

CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era

Model ReleasesDGX agent

arXiv:2602.23452v2 Announce Type: replace Abstract: Scientific research relies on accurate citation for attribution and integrity, yet large language models (LLMs) introduce a new risk: fabricated ref

‘Code is cheap. Mistakes are expensive’: Appian gives vibe coding an enterprise reality check

ApplicationsDGX agent

The rise of vibe coding is pushing software teams to build faster than ever before, but enterprise buyers demand systems that can survive audits and operational risk. The result is a new kind of softw

Credal Concept Bottleneck Models for Epistemic-Aleatoric Uncertainty Decomposition

ResearchDGX agent

arXiv:2604.24170v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) predict through human-interpretable concepts, but they typically output point concept probabilities that conflate epist

CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks

Model ReleasesDGX agent

arXiv:2510.17687v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) achieve strong reasoning and perception capabilities but are increasingly vulnerable to jailbreak att

CyberCane: Neuro-Symbolic RAG for Privacy-Preserving Phishing Detection with Formal Ontology Reasoning

ApplicationsDGX agent

arXiv:2604.23563v1 Announce Type: cross Abstract: Privacy-critical domains require phishing detection systems that satisfy contradictory constraints: near-zero false positives to prevent workflow disr

Domain-Filtered Knowledge Graphs from Sparse Autoencoder Features

Local AiDGX agent

arXiv:2604.23829v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) extract millions of interpretable features from a language model, but flat feature inventories aren't very useful on their ow

Fragmented AI policy threatens US leadership as government scrambles to keep pace

SafetyDGX agent

AI policy fragmentation is emerging as a critical risk for Washington, and without a federal standard, a patchwork of conflicting state-level rules threatens to undermine American competitiveness. Tha

From months to minutes: Building real-time clinical data pipelines with natural language

IndustryDGX agent

This Databricks blog post discusses how natural language processing can accelerate clinical data pipeline development, reducing implementation time from months to minutes. It covers techniques for ext

From Rights to Rites: Expectations Management in Smart-Home AI

ResearchDGX agent

arXiv:2604.23635v1 Announce Type: cross Abstract: Domestic voice assistants and smart-home devices are increasingly embedded in everyday routines, yet their ethics are often treated as an afterthought

Gartner: US states issued $3.45B in privacy-related fines to companies in 2025, a total larger than the last five years combined, driven by new privacy laws (Derek B. Johnson/CyberScoop)

IndustryDGX agent

Derek B. Johnson / CyberScoop: Gartner: US states issued $3.45B in privacy-related fines to companies in 2025, a total larger than the last five years combined, driven by new privacy laws — The increa

Interpretable Physics-Informed Load Forecasting for U.S. Grid Resilience: SHAP-Guided Ensemble Validation in Hybrid Deep Learning Under Extreme Weather

ResearchDGX agent

arXiv:2604.23500v1 Announce Type: cross Abstract: Accurate short-term electricity load forecasting is a cornerstone of U.S. grid reliability; however, prevailing deep learning models remain opaque, li

IntrAgent: An LLM Agent for Content-Grounded Information Retrieval through Literature Review

Model ReleasesDGX agent

arXiv:2604.22861v1 Announce Type: cross Abstract: Scientific research relies on accurate information retrieval from literature to support analytical decisions. In this work, we introduce a new task, I

KLong: Training LLM Agent for Extremely Long-horizon Tasks

Model ReleasesDGX agent

arXiv:2602.17547v3 Announce Type: replace Abstract: This paper introduces KLong, an open-source LLM agent trained to solve extremely long-horizon tasks. The principle is to first cold-start the model

Machine Learning for Network Attacks Classification and Statistical Evaluation of Adversarial Learning Methodologies for Synthetic Data Generation

ResearchDGX agent

arXiv:2603.17717v3 Announce Type: replace-cross Abstract: Supervised detection of network attacks has always been a critical part of network intrusion detection systems (NIDS). Nowadays, in a pivotal

MEMCoder: Multi-dimensional Evolving Memory for Private-Library-Oriented Code Generation

Model ReleasesDGX agent

arXiv:2604.24222v1 Announce Type: cross Abstract: Large Language Models (LLMs) excel at general code generation, but their performance drops sharply in enterprise settings that rely on internal privat

Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs

Model ReleasesDGX agent

arXiv:2509.04802v3 Announce Type: replace Abstract: As large language models increasingly deployed into agentic systems, existing methods face critical gaps in observing, assessing, and mitigating dep

NeuroClaw Technical Report

Model ReleasesDGX agent

arXiv:2604.24696v1 Announce Type: new Abstract: Agentic artificial intelligence systems promise to accelerate scientific workflows, but neuroimaging poses unique challenges: heterogeneous modalities (

new in ml-intern: you can now actually see what's going on inside added native metric logging + trackio integration. every training run the …

Model ReleasesDGX agent

new in ml-intern: you can now actually see what's going on inside added native metric logging + trackio integration. every training run the agent kicks off now has live curves you can watch in real ti

OmniSch: A Multimodal PCB Schematic Benchmark For Structured Diagram Visual Reasoning

Model ReleasesDGX agent

arXiv:2604.00270v2 Announce Type: replace Abstract: Recent large multimodal models (LMMs) have made rapid progress in visual grounding, document understanding, and diagram reasoning tasks. However, th

OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents

SafetyDGX agent

arXiv:2604.24348v1 Announce Type: new Abstract: The evolution of Multimodal Large Language Models (MLLMs) has shifted the focus from text generation to active behavioral execution, particularly via OS

Phenom adds Plum psychometric science to its agentic AI hiring stack

AgentsDGX agent

Artificial intelligence-based human resources company Phenom People Inc. announced today that it has acquired Plum.io Inc., a psychometric-based talent assessments company that measures the durable sk

PhysCodeBench: Benchmarking Physics-Aware Symbolic Simulation of 3D Scenes via Self-Corrective Multi-Agent Refinement

Model ReleasesDGX agent

arXiv:2604.23580v1 Announce Type: cross Abstract: Physics-aware symbolic simulation of 3D scenes is critical for robotics, embodied AI, and scientific computing, requiring models to understand natural

SAS expands agentic AI and governance capabilities with broad platform updates

AgentsDGX agent

SAS Institute Inc. is using its SAS Innovate conference this week to introduce a sweeping set of product updates aimed at helping enterprises operationalize artificial intelligence with stronger gover

Scalable LLM-based Coding of Dialogue in Healthcare Simulation: Balancing Coding Performance, Processing Time, and Environmental Impact

ApplicationsDGX agent

arXiv:2604.23255v1 Announce Type: cross Abstract: Research shows that dialogue, the interactive process through which participants articulate their thinking, plays a central role in constructing share

Secure On-Premise Deployment of Open-Weights Large Language Models in Radiology: An Isolation-First Architecture with Prospective Pilot Evaluation

Model ReleasesDGX agent

arXiv:2604.22768v1 Announce Type: cross Abstract: Purpose: To design, implement, evaluate, and report on the regulatory requirements of a self-hosted LLM infrastructure for radiology adhering to the p

Self-Admitted Technical Debt Detection Approaches: A Decade Systematic Review

ResearchDGX agent

arXiv:2312.15020v4 Announce Type: replace-cross Abstract: Technical debt (TD) refers to the long-term costs associated with suboptimal design or code decisions in software development, often made to m

ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction

Model ReleasesDGX agent

arXiv:2604.23813v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable performance in Visually Rich Document Understanding (VRDU) tasks, but their capabili

since we have asks in one place, it makes it easier to analyze where founders are asking us for help we found 75% of asks are around access …

AgentsDGX agent

since we have asks in one place, it makes it easier to analyze where founders are asking us for help we found 75% of asks are around access to individuals or types of poeple next step here is analyze

SPLIT: Separating Physical-Contact via Latent Arithmetic in Image-Based Tactile Sensors

ResearchDGX agent

arXiv:2604.24449v1 Announce Type: cross Abstract: Training machine learning models for robotic tactile sensing requires vast amounts of data, yet obtaining realistic interaction data remains a challen

SWE-QA: Can Language Models Answer Repository-level Code Questions?

Model ReleasesDGX agent

arXiv:2509.14635v2 Announce Type: replace Abstract: Understanding and reasoning about entire software repositories is an essential capability for intelligent software engineering tools. While existing

SycoPhantasy: Quantifying Sycophancy and Hallucination in Small Open Weight VLMs for Vision-Language Scoring of Fantasy Characters

Model ReleasesDGX agent

arXiv:2604.24346v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed as evaluators in tasks requiring nuanced image understanding, yet their reliability in scoring

TeachMaster: Generative Teaching via Code

AgentsDGX agent

arXiv:2601.04204v2 Announce Type: replace-cross Abstract: The scalability of high-quality online education is hindered by the high costs and slow cycles of manual content creation. Despite advancement

The Last Human-Written Paper: Agent-Native Research Artifacts

AgentsDGX agent

arXiv:2604.24658v1 Announce Type: new Abstract: Scientific publication compresses a branching, iterative research process into a linear narrative, discarding the majority of what was discovered along

Time-Series Forecasting in Safety-Critical Environments: An EU-AI-Act-Compliant Open-Source Package / Zeitreihenprognose in sicherheitskritischen Umgebungen: Ein KI-VO-konformes Open-Source-Paket

SafetyDGX agent

arXiv:2604.23859v1 Announce Type: new Abstract: With spotforecast2-safe we present an integrated Compliance-by-Design approach to Python-based point forecasting of time series in safety-critical envir

Towards Lawful Autonomous Driving: Deriving Scenario-Aware Driving Requirements from Traffic Laws and Regulations

AgentsDGX agent

arXiv:2604.24562v1 Announce Type: new Abstract: Driving in compliance with traffic laws and regulations is a basic requirement for human drivers, yet autonomous vehicles (AVs) can violate these requir

Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills

Model ReleasesDGX agent

arXiv:2603.25158v4 Announce Type: replace Abstract: Equipping Large Language Model (LLM) agents with domain-specific skills is critical for tackling complex tasks. Yet, manual authoring creates a seve

Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines

SafetyDGX agent

arXiv:2604.23001v1 Announce Type: cross Abstract: Despite remarkable progress in Vision--Language--Action (VLA) models, a central bottleneck remains underexamined: the data infrastructure that underli

Weakly Supervised Multicenter Nancy Index Scoring in Ulcerative Colitis Using Foundation Models

ResearchDGX agent

arXiv:2604.23706v1 Announce Type: new Abstract: Histologic assessment of ulcerative colitis (UC) activity is an important endpoint in clinical trials and routine care, but manual grading with indices

← Previous
1…7273747576…83
Next →