AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “research”

GridTimelineEvolution
25,874 results
12 Aug 2026

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

SafetyDGX agent

arXiv:2608.10621v1 Announce Type: new Abstract: Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formula

Recovering Wasted Compute in Autoresearch Agents

AgentsDGX agent

arXiv:2608.10424v1 Announce Type: new Abstract: A slew of recent works develop agents for solving research problems end-to-end, a paradigm increasingly referred to as autoresearch. Such agents have in

Seq2Synth: Benchmarking Temporal Fidelity in Synthetic Sequential Tabular Data

Model ReleasesDGX agent

arXiv:2607.15606v2 Announce Type: replace Abstract: Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing and data-driven research, but evaluating their fidelity

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Toward a Theory of Value in AI Alignment

SafetyDGX agent

arXiv:2608.10327v1 Announce Type: new Abstract: Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms

11 Aug 2026

Accurate but Natural? Diagnosing Grammatical and Idiomatic Gaps in Japanese EFL Writing

ApplicationsDGX agent

arXiv:2608.09289v1 Announce Type: new Abstract: Second language writing research distinguishes grammatical accuracy from native-like idiomaticity, yet automated writing evaluation often conflates thes

Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy

SafetyDGX agent

arXiv:2608.09857v1 Announce Type: cross Abstract: Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on

Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models

Model ReleasesDGX agent

arXiv:2608.08822v1 Announce Type: new Abstract: Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and b

Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks

Model ReleasesDGX agent

arXiv:2603.00077v3 Announce Type: replace-cross Abstract: Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduc

Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production

ApplicationsDGX agent

arXiv:2608.09045v1 Announce Type: cross Abstract: Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such as isolate

Efficient Human-Contact Representation for Human-Scene Interaction

Model ReleasesDGX agent

arXiv:2608.09388v1 Announce Type: new Abstract: Human-scene interaction is an active research topic with several industrial applications in virtual reality, gaming, robotics, and surveillance. Despite

Introducing 𝗘𝘅𝘁𝗿𝗮𝗰𝘁𝗕𝗲𝗻𝗰𝗵: the most comprehensive benchmark for information extraction from complex enterprise documents. Our app…

Model ReleasesDGX agent

Introducing 𝗘𝘅𝘁𝗿𝗮𝗰𝘁𝗕𝗲𝗻𝗰𝗵: the most comprehensive benchmark for information extraction from complex enterprise documents. Our applied research team tested: 14 systems — frontier VLMs, coding agents, ex

Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family

SafetyDGX agent

arXiv:2604.05971v2 Announce Type: replace-cross Abstract: Recent research has shown that contrastive vision-language models such as CLIP often lack fine-grained understanding of visual content. While

Parameter Exploration for RLVR via Variational Learning

Model ReleasesDGX agent

arXiv:2608.09805v1 Announce Type: cross Abstract: Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an importan

PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary

Model ReleasesDGX agent

arXiv:2608.08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically frame

SPD Learn: A Geometric Deep Learning Python Library for Neural Decoding Through Trivialization

TutorialsDGX agent

arXiv:2602.22895v2 Announce Type: replace-cross Abstract: Implementations of symmetric positive definite (SPD) matrix-based neural networks for neural decoding remain fragmented across research codeba

The Cell Must Go On: Agar.io for Continual Reinforcement Learning

Model ReleasesDGX agent

arXiv:2505.18347v3 Announce Type: replace-cross Abstract: Continual reinforcement learning (RL) concerns agents that are expected to learn continually, rather than converge to a policy that is then fi

Towards Adaptive Super-Resolution and Quality Assessment via Test-Time Adaptation

ApplicationsDGX agent

arXiv:2608.08508v1 Announce Type: new Abstract: This paper presents doctoral research on adaptive video super-resolution and perceptual quality modeling under real-world conditions. Existing video sup

When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery

Model ReleasesDGX agent

arXiv:2608.08132v1 Announce Type: new Abstract: Reconstruction of 3D objects from a single image is a challenging research problem in computer vision. The key challenge is the lack of critical informa

Zero-shot 2D Grounding with Novel Affordance Types

Model ReleasesDGX agent

arXiv:2608.08929v1 Announce Type: new Abstract: 2D affordance grounding aims to locate the region of an object that a human can interact with. Existing research focuses on recognizing affordance types

10 Aug 2026

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

Model ReleasesDGX agent

arXiv:2608.07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate t

How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots

SafetyDGX agent

arXiv:2608.06898v1 Announce Type: cross Abstract: Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public l

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents

SafetyDGX agent

arXiv:2608.06663v1 Announce Type: new Abstract: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour

9 Aug 2026

CyberKimi just dropped strong results on one of ExploitBench’s hardest V8 bugs , points away from Mythos

Model ReleasesDGX agent

Hey everyone ! Quick share from the cyber + local LLM side of things that I found interesting. During this week’s hacker summer camp, an AI researcher and reverse malware engineer veteran 'lordx64' on

KLQ: Training-free measured rotation quantization. Beats all training-free rotation-based quantization methods on W4A4KV4-bits. Llama 3.2 1B KLQ-quantized beats SpinQuant and gets close to ReSpinQuant without GPTQ/LDLQ rounding.

Model ReleasesDGX agent

First of all, I'm not a lab, this was a solo summer research project that finally culminated into the github repo and the writeup. The repo includes a much deeper dive with methods, findings about qua

SQLite compressed text-history prototypes

Model ReleasesDGX agent

Research: SQLite compressed text-history prototypes I'm perennially interested in options for storing revision histories in relational databases. While out on a dog walk I had a new idea: how about ta

8 Aug 2026

Quick survey (2 min) on trust in hardware specs for open-source models

Local AiDGX agent

Hi everyone, I'm a systems analysis student researching a problem a lot of you probably know well: how much you actually trust the published VRAM/RAM requirements for open-source models before trying

We analyzed DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. A DeepSeek-first cascade with test-suite verification solved MORE tasks than Luna…

Model ReleasesDGX agent

Researchers from TogetherAI analyzed DeepSeek V4 Flash and GPT‑5.6 Luna on the DeepSWE benchmark. The study found that employing a DeepSeek‑first cascade with test‑suite verification solved more tasks

7 Aug 2026

Agentic Software Issue Resolution with Large Language Models: A Survey

AgentsDGX agent

arXiv:2512.22256v2 Announce Type: replace-cross Abstract: Software issue resolution aims to address real-world issues in software repositories based on natural language descriptions provided by users,

AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic Simulations

Model ReleasesDGX agent

arXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr

Deep Generalised Mixed Models: a Novel Neural Network Structure for Analysing Hierarchical Data

Model ReleasesDGX agent

arXiv:2608.05930v1 Announce Type: cross Abstract: The experience sampling method (ESM) is a longitudinal research design where participants report their thoughts, emotional states and behaviours multi

How TReNDS automates root-cause analysis with Amazon Bedrock

AgentsDGX agent

TReNDS, a research center at Georgia State University, built an agentic AI pipeline on Amazon Bedrock and the open-source Strands Agents SDK that automatically investigates production errors in real t

Measuring and Detecting Harmful AI Sycophancy

TutorialsDGX agent

arXiv:2608.05624v1 Announce Type: new Abstract: Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful. This pa

Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges

SafetyDGX agent

arXiv:2405.15604v4 Announce Type: replace Abstract: Text generation has become more accessible than ever, and the growing interest in these systems, especially those using large language models, has s

TLNM: Externally Validated Tooth Detection, Numbering and Segmentation from Smartphone Photographs Using Mask R-CNN

Local AiDGX agent

arXiv:2608.06275v1 Announce Type: new Abstract: Oral health issues affect billions globally, but the cost and limited access to professional dental care hinder preventive oral healthcare. Research rel

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

AgentsDGX agent

arXiv:2608.06366v1 Announce Type: new Abstract: Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists' workload

Unifying Structured and Unstructured Data Insights with BQ Search Innovations

Model ReleasesDGX agent

Modern enterprises possess a vast amount of unstructured data, yet they frequently encounter significant challenges in managing and extracting value from it. Historically, unlocking the insights hidde

6 Aug 2026

ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation

AgentsDGX agent

arXiv:2608.05026v1 Announce Type: cross Abstract: High-quality annotation of artworks is essential for computational art research, yet extracting implicit semantics remains challenging due to the reli

Eigenius: A Typed Knowledge-Graph DBMS with Epistemic Stratification and Institution-Mediated Reasoning

AgentsDGX agent

arXiv:2608.04457v1 Announce Type: cross Abstract: As 'AI Scientists' emerge to drive research via the Model Context Protocol (MCP), systems relying on ephemeral scripts will fail. The sheer scale of s

FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation

Model ReleasesDGX agent

arXiv:2608.04374v1 Announce Type: cross Abstract: Large language models can produce fluent financial analysis, but fluency alone does not establish whether a report is suitable for institutional deliv

IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath)

Model ReleasesDGX agent

arXiv:2608.04703v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for question answering, education, and research, including in religious and cultural domains where an

K-EXAONE 2.0 Technical Report

Model ReleasesDGX agent

arXiv:2608.04505v1 Announce Type: new Abstract: This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research as a step in our effort toward glo

NSF-HRPT: Neural Semantic Field meets Hierarchical Risk Perception Tree for Safety-Critical Scenario Assessment

SafetyDGX agent

arXiv:2608.04776v1 Announce Type: new Abstract: The ability to accurately assess and anticipate risks in safety-critical scenarios is crucial for autonomous driving systems. While existing research ha

Promptable Animal Pose Tracking Across Species

Local AiDGX agent

arXiv:2608.04995v1 Announce Type: new Abstract: Animal pose estimation and tracking is important for wildlife monitoring and conservation research, and with limited expert time for labelling automated

SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models

Model ReleasesDGX agent

arXiv:2608.04975v1 Announce Type: cross Abstract: SciCode is the standard measure of the scientific-coding ability of language models: research-level problems that demand both frontier scientific theo

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

Model ReleasesDGX agent

arXiv:2608.04317v1 Announce Type: cross Abstract: Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost

5 Aug 2026

A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models

SafetyDGX agent

arXiv:2608.02684v1 Announce Type: cross Abstract: Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in

Accelerating Human-Aware Robot Trajectory Generation via Diffusion and Consistency Distillation

SafetyDGX agent

arXiv:2608.03159v1 Announce Type: new Abstract: This research proposes a constrained motion planning framework for robot manipulators in human-robot interaction (HRI). For a non-redundant manipulator

Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls

Model ReleasesDGX agent

arXiv:2608.03071v1 Announce Type: new Abstract: Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting the right tool

Harness choice is a big deal. So much room to advance and improve results across the board with agent harnesses. Great paper highlighting th…

Model ReleasesDGX agent

Harness choice is a big deal. So much room to advance and improve results across the board with agent harnesses. Great paper highlighting this. New research releases DataSpace, a benchmark where data

Scaling agentic AI: How UiPath built its high-performance GPU platform on AI Hypercomputer

Model ReleasesDGX agent

As a market leader in enterprise agentic automation and business orchestration, UiPath is helping to pioneer an industry shift toward agentic AI. With it, the company is deploying autonomous agents to

Self-Supervised Uncalibrated Multi-View Video Anonymization in the Operating Room

Local AiDGX agent

arXiv:2602.02850v3 Announce Type: replace Abstract: Privacy preservation is a prerequisite for using video data in Operating Room (OR) research. Effective anonymization relies on the exhaustive locali

Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI…

SafetyDGX agent

Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI model framework. I'm not surprised at all, and it's actuall

StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision

Model ReleasesDGX agent

arXiv:2603.29368v2 Announce Type: replace Abstract: Driven by the advancement of 3D devices, stereo vision tasks including stereo matching and stereo conversion have emerged as a critical research fro

When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Coupling Diagnostic for Machine Collectives

Model ReleasesDGX agent

arXiv:2608.03722v1 Announce Type: new Abstract: Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capa

4 Aug 2026

MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs

Model ReleasesDGX agent

arXiv:2608.02520v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for health-related advice. Existing research measures their safety with static questions rather than

one thing i appreciate about silico is that it's a deeply humanist product. we designed silico to keep you in the experimental loop -- more …

Model ReleasesDGX agent

one thing i appreciate about silico is that it's a deeply humanist product. we designed silico to keep you in the experimental loop -- more observable, easier to steer, easier to understand we want to

Passing Coarse Marginal Checks Can Be Cheap: Persona Mixtures and Imprecise Treatment-Response Estimates in an LLM Persona Panel

Model ReleasesDGX agent

arXiv:2608.00979v1 Announce Type: cross Abstract: Large language models are increasingly used as synthetic research participants and are often validated by whether their marginal responses resemble hu

RadYOLO: Computationally Efficient 3D Object Detection and Segmentation in CT and MRI

Local AiDGX agent

arXiv:2608.00508v1 Announce Type: new Abstract: Object detection and segmentation in three-dimensional medical images is a very active area of research. However, most proposed deep learning models car

3 Aug 2026

EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment

Model ReleasesDGX agent

arXiv:2604.08342v2 Announce Type: replace Abstract: Long context egocentric video understanding has recently attracted significant research attention, with augmented reality (AR) highlighted as one of

Geographically Weighted Surrogate Models for Rapid Small-Area Chronic Disease Estimation

Model ReleasesDGX agent

arXiv:2607.28655v1 Announce Type: cross Abstract: Small-area estimation (SAE) enables researchers and policymakers to identify spatial disparities in health outcomes, but survey-based SAE products car

← Previous
1…347348349350351…432
Next →