AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “research”

GridTimelineEvolution
25,425 results
15 Apr 2026

Siamese Foundation Models for Crystal Structure Prediction

TutorialsDGX agent

arXiv:2503.10471v2 Announce Type: replace-cross Abstract: Predicting crystal structures from chemical compositions is a fundamental challenge in materials discovery, complicated by complex 3D geometri

SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks

Model ReleasesDGX agent

arXiv:2506.14512v4 Announce Type: replace Abstract: Large Language Models (LLMs) have undergone rapid progress, largely attributed to reinforcement learning on complex reasoning tasks. In contrast, wh

Thermodynamic Liquid Manifold Networks: Physics-Bounded Deep Learning for Solar Forecasting in Autonomous Off-Grid Microgrids

AgentsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.11909v1 Announce Type: cross Abstract: The stable operation of autonomous off-grid photovoltaic systems requires solar forecasting algorithms that respect atmospheric thermodynamics. Contem

This sub is mostly screenshots of ChatGPT being wrong. Meanwhile people on RunLobster (OpenClaw) are quietly running real businesses on AI.

IndustryDGX agent

This Reddit post contrasts the r/ChatGPT community's tendency to focus on AI failures and viral screenshots with a more pragmatic user base building real workflows on RunLobster, a managed cloud platf

TRUST Agents: A Collaborative Multi-Agent Framework for Fake News Detection, Explainable Verification, and Logic-Aware Claim Reasoning

Model ReleasesDGX agent

arXiv:2604.12184v1 Announce Type: new Abstract: TRUST Agents is a collaborative multi-agent framework for explainable fact verification and fake news detection. Rather than treating verification as a

Uncertainty Quantification on Graph Learning: A Survey

ApplicationsDGX agent

arXiv:2404.14642v4 Announce Type: replace Abstract: Graphical models have demonstrated their exceptional capabilities across numerous applications. However, their performance, confidence, and trustwor

WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces

SafetyDGX agent

arXiv:2603.05295v3 Announce Type: replace Abstract: We introduce WebChain, the largest open-source dataset of human-annotated trajectories on real-world websites, designed to accelerate reproducible r

'We're going to a world where we're building systems that will be smart to us not like Einstein is to an average person, but like humans are to mice or ants'

IndustryDGX agent

This Reddit post on r/ChatGPT discusses a striking quote about the anticipated intelligence gap between future AI systems and humans — comparing it not to the relatively modest difference between a ge

Why Did Apple Fall: Evaluating Curiosity in Large Language Models

TutorialsDGX agent

arXiv:2510.20635v2 Announce Type: replace-cross Abstract: Curiosity serves as a pivotal conduit for human beings to discover and learn new knowledge. Recent advancements of large language models (LLMs

14 Apr 2026

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness

Model ReleasesDGX agent

arXiv:2604.09564v1 Announce Type: cross Abstract: We present ACE-Bench (Azure SDK Coding Evaluation Benchmark), an execution-free benchmark that provides fast, reproducible pass or fail signals for wh

Agentic Exploration of PDE Spaces using Latent Foundation Models for Parameterized Simulations

Model ReleasesDGX agent

arXiv:2604.09584v1 Announce Type: new Abstract: Flow physics and more broadly physical phenomena governed by partial differential equations (PDEs), are inherently continuous, high-dimensional and ofte

AI Integrity: A New Paradigm for Verifiable AI Governance

SafetyDGX agent

arXiv:2604.11065v1 Announce Type: new Abstract: AI systems increasingly shape high-stakes decisions in healthcare, law, defense, and education, yet existing governance paradigms -- AI Ethics, AI Safet

AI sycophancy is 41% worse on philosophy than math - and varies by who's asking, new study finds

IndustryDGX agent

A study on AI sycophancy finds that the phenomenon varies significantly by domain and user context — with subjective, opinion-based fields like philosophy showing substantially higher rates of AI agre

[AINews] Top Local Models List - April 2026

ToolsDGX agent

a quiet day lets us check in on the local models scene

Assessing Model-Agnostic XAI Methods against EU AI Act Explainability Requirements

SafetyDGX agent

arXiv:2604.09628v1 Announce Type: cross Abstract: Explainable AI (XAI) has evolved in response to expectations and regulations, such as the EU AI Act, which introduces regulatory requirements on AI-po

Automatic Uncertainty-Aware Synthetic Data Bootstrapping for Historical Map Segmentation

ApplicationsDGX agent

arXiv:2511.15875v2 Announce Type: replace Abstract: The automated analysis of historical documents, particularly maps, has drastically benefited from advances in deep learning and its success across v

BEM: Training-Free Background Embedding Memory for False-Positive Suppression in Real-Time Fixed-Background Camera

ApplicationsDGX agent

arXiv:2604.11714v1 Announce Type: new Abstract: Pretrained detectors perform well on benchmarks but often suffer performance degradation in real-world deployments due to distribution gaps between trai

Beyond Fixed False Discovery Rates: Post-Hoc Conformal Selection with E-Variables

ApplicationsDGX agent

arXiv:2604.11305v1 Announce Type: new Abstract: Conformal selection (CS) uses calibration data to identify test inputs whose unobserved outcomes are likely to satisfy a pre-specified minimal quality r

Beyond Message Passing: A Semantic View of Agent Communication Protocols

SafetyDGX agent

arXiv:2604.02369v3 Announce Type: replace-cross Abstract: Agent communication protocols are becoming critical infrastructure for large language model (LLM) systems that must use tools, coordinate with

Big lab leaks

IndustryDGX agent

AI development is shifting toward fully integrated, agent-enabled applications, with major platforms like Anthropic's Claude and OpenAI

Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptation

ApplicationsDGX agent

arXiv:2604.10950v1 Announce Type: new Abstract: Fully supervised Video Semantic Segmentation (VSS) relies heavily on densely annotated video data, limiting practical applicability. Alternatively, appl

C-ReD: A Comprehensive Chinese Benchmark for AI-Generated Text Detection Derived from Real-World Prompts

Model ReleasesDGX agent

arXiv:2604.11796v1 Announce Type: cross Abstract: Recently, large language models (LLMs) are capable of generating highly fluent textual content. While they offer significant convenience to humans, th

Can Large Language Models Infer Causal Relationships from Real-World Text?

Model ReleasesDGX agent

arXiv:2505.18931v4 Announce Type: replace Abstract: Understanding and inferring causal relationships from texts is a core aspect of human cognition and is essential for advancing large language models

CapyMOA: Efficient Machine Learning for Data Streams and Online Continual Learning in Python

TutorialsDGX agent

arXiv:2502.07432v2 Announce Type: replace Abstract: CapyMOA is an open-source Python library for efficient machine learning on data streams and online continual learning. It provides a structured fram

Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

Model ReleasesDGX agent

arXiv:2503.21380v3 Announce Type: replace Abstract: The rapid advancement of large reasoning models has saturated existing math benchmarks, underscoring the urgent need for more challenging evaluation

ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents

AgentsDGX agent

arXiv:2509.22830v3 Announce Type: replace Abstract: The growing deployment of large language model (LLM) based agents that interact with external environments has created new attack surfaces for adver

Comparative Analysis of Large Language Models in Healthcare

Model ReleasesDGX agent

arXiv:2604.10316v1 Announce Type: new Abstract: Background: Large Language Models (LLMs) are transforming artificial intelligence applications in healthcare due to their ability to understand, generat

Cybersecurity Looks Like Proof of Work Now

Model ReleasesDGX agent

Cybersecurity Looks Like Proof of Work Now The UK's AI Safety Institute recently published Our evaluation of Claude Mythos Preview’s cyber capabilities, their own independent analysis of Claude Mythos

Deep Learning for Sequential Decision Making under Uncertainty: Foundations, Frameworks, and Frontiers

AgentsDGX agent

arXiv:2604.11507v1 Announce Type: cross Abstract: Artificial intelligence (AI) is moving increasingly beyond prediction to support decisions in complex, uncertain, and dynamic environments. This shift

DiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary Domain

Model ReleasesDGX agent

arXiv:2604.10425v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) have revolutionized general visual understanding. However, their application in the food domain rem

Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky

Model ReleasesDGX agent

arXiv:2507.03336v4 Announce Type: replace Abstract: Large language models (LLMs) are increasingly tasked with invoking enterprise APIs, yet they routinely falter when near-duplicate tools vie for the

Distributionally Robust PAC-Bayesian Control

SafetyDGX agent

arXiv:2604.10588v1 Announce Type: new Abstract: We present a distributionally robust PAC-Bayesian framework for certifying the performance of learning-based finite-horizon controllers. While existing

Doc-PP: Document Policy Preservation Benchmark for Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2601.03926v2 Announce Type: replace Abstract: The deployment of Large Vision-Language Models (LVLMs) for real-world document question answering is often constrained by dynamic, user-defined poli

DocRevive: A Unified Pipeline for Document Text Restoration

Model ReleasesDGX agent

arXiv:2604.10077v1 Announce Type: new Abstract: In Document Understanding, the challenge of reconstructing damaged, occluded, or incomplete text remains a critical yet unexplored problem. Subsequent d

Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection

ApplicationsDGX agent

arXiv:2604.09920v1 Announce Type: new Abstract: Vision foundation models (VFMs) offer the promise of zero-shot object detection without task-specific training data, yet their performance in complex ag

Domain-Specific Data Generation Framework for RAG Adaptation

ApplicationsDGX agent

arXiv:2510.11217v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) combines the language understanding and reasoning power of large language models (LLMs) with external ret

DoSReMC: Domain Shift Resilient Mammography Classification using Batch Normalization Adaptation

ApplicationsDGX agent

arXiv:2508.15452v3 Announce Type: replace-cross Abstract: Numerous deep learning-based solutions have been developed for the automatic recognition of breast cancer using mammography images. However, t

Dual-Margin Embedding for Fine-Grained Long-Tailed Plant Taxonomy

ApplicationsDGX agent

arXiv:2512.18994v2 Announce Type: replace Abstract: Taxonomic classification of ecological families, genera, and species underpins biodiversity monitoring and conservation. Existing computer vision me

EDFNet: Early Fusion of Edge and Depth for Thin-Obstacle Segmentation in UAV Navigation

AgentsDGX agent

arXiv:2604.09694v1 Announce Type: new Abstract: Autonomous Unmanned Aerial Vehicles (UAVs) must reliably detect thin obstacles such as wires, poles, and branches to navigate safely in real-world envir

Empowering Video Translation using Multimodal Large Language Models

SafetyDGX agent

arXiv:2604.11283v1 Announce Type: new Abstract: Recent developments in video translation have further enhanced cross-lingual access to video content, with multimodal large language models (MLLMs) play

Ernie Image Turbo is Capable of ...

Local AiDGX agent

ERNIE Image Turbo is an open text-to-image generation model developed by Baidu's ERNIE-Image team, serving as the distilled release of the full ERNIE-Image model and built on a single-stream Diffusion

Evaluating Cooperation in LLM Social Groups through Elected Leadership

AgentsDGX agent

arXiv:2604.11721v1 Announce Type: cross Abstract: Governing common-pool resources requires agents to develop enduring strategies through cooperation and self-governance to avoid collective failure. Wh

Examining EAP Students' AI Disclosure Intention: A Cognition-Affect-Conation Perspective

SafetyDGX agent

arXiv:2604.10991v1 Announce Type: cross Abstract: The growing use of generative artificial intelligence (AI) in academic writing has raised increasing concerns regarding transparency and academic inte

Explainability and Certification of AI-Generated Educational Assessments

SafetyDGX agent

arXiv:2604.09622v1 Announce Type: cross Abstract: The rapid adoption of generative artificial intelligence (AI) in educational assessment has created new opportunities for scalable item creation, pers

Exploring the impact of fairness-aware criteria in AutoML

SafetyDGX agent

arXiv:2604.10224v1 Announce Type: cross Abstract: Machine Learning (ML) systems are increasingly used to support decision-making processes that affect individuals. However, these systems often rely on

Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation

TutorialsDGX agent

arXiv:2510.10925v2 Announce Type: replace-cross Abstract: Training student models on synthetic data generated by strong teacher models is a promising way to distilling the capabilities of teachers. Ho

From GPT-3 to GPT-5: Mapping their capabilities, scope, limitations, and consequences

Model ReleasesDGX agent

arXiv:2604.10332v1 Announce Type: new Abstract: We present the progress of the GPT family from GPT-3 through GPT-3.5, GPT-4, GPT-4 Turbo, GPT-4o, GPT-4.1, and the GPT-5 family. Our work is comparative

From Helpful to Trustworthy: LLM Agents for Pair Programming

AgentsDGX agent

arXiv:2604.10300v1 Announce Type: cross Abstract: LLM-based coding agents are increasingly used to generate code, tests, and documentation. Still, their outputs can be plausible yet misaligned with de

GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents

AgentsDGX agent

arXiv:2603.24329v2 Announce Type: replace-cross Abstract: Multimodal LLMs are increasingly deployed as perceptual backbones for autonomous agents in 3D environments, from robotics to virtual worlds. T

General-purpose LLMs as Models of Human Driver Behavior: The Case of Simplified Merging

Model ReleasesDGX agent

arXiv:2604.09609v1 Announce Type: new Abstract: Human behavior models are essential as behavior references and for simulating human agents in virtual safety assessment of automated vehicles (AVs), yet

GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2505.17022v2 Announce Type: replace-cross Abstract: Visual generation models have made remarkable progress in creating realistic images from text prompts, yet struggle with complex prompts that

How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities

Model ReleasesDGX agent

arXiv:2603.02578v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly deployed in socially sensitive domains, yet their unpredictable behaviors, ranging from misalign

I gave an AI agent acess to my calendar and email for two weeks and here is what i actually learned.

AgentsDGX agent

A Reddit user on r/ChatGPT conducted a real-world, two-week personal experiment granting an AI agent direct access to their calendar and email, documenting practical outcomes and honest takeaways. The

Influencing Humans to Conform to Preference Models for RLHF

SafetyDGX agent

arXiv:2501.06416v3 Announce Type: replace-cross Abstract: Designing a reinforcement learning from human feedback (RLHF) algorithm to approximate a human's unobservable reward function requires assumin

Integrating SAINT with Tree-Based Models: A Case Study in Employee Attrition Prediction

ApplicationsDGX agent

arXiv:2604.10337v1 Announce Type: new Abstract: Employee attrition presents a major challenge for organizations, increasing costs and reducing productivity. Predicting attrition accurately enables pro

Intelligent bear deterrence system based on computer vision: Reducing human bear conflicts in remote areas

Local AiDGX agent

arXiv:2503.23178v2 Announce Type: replace Abstract: Conflicts between humans and bears on the Tibetan Plateau present substantial threats to local communities and hinder wildlife preservation initiati

Interactive Interface For Semantic Segmentation Dataset Synthesis

ApplicationsDGX agent

arXiv:2506.23470v2 Announce Type: replace Abstract: The rapid advancement of AI and computer vision has significantly increased the demand for high-quality annotated datasets, particularly for semanti

Interesting: 'Currently, 38% of Americans live within 5 miles of at least one operational data center... Living near a data center doesn’t h…

ApplicationsDGX agent

Interesting: 'Currently, 38% of Americans live within 5 miles of at least one operational data center... Living near a data center doesn’t have much of an effect on public opinion about the facilities

Is there a workflow to relight videos with perfect pixel-level alignment?

SafetyDGX agent

This r/StableDiffusion thread discusses community-driven approaches to relighting videos using Stable Diffusion-based tools, with a focus on the challenge of maintaining pixel-perfect alignment betwee

ks-pret-5m: a 5 million word, 12 million token kashmiri pretraining dataset

Model ReleasesDGX agent

arXiv:2604.11066v1 Announce Type: new Abstract: We present KS-PRET-5M, the largest publicly available pretraining dataset for the Kashmiri language, comprising 5,090,244 (5.09M) words, 27,692,959 (27.

← Previous
1…418419420421422…424
Next →