AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “research”

GridTimelineEvolution
25,639 results
3 Jul 2026

Provably Finding a Hidden Dense Submatrix among Many Planted Dense Submatrices via Convex Programming

ApplicationsDGX agent

arXiv:2601.03946v3 Announce Type: replace-cross Abstract: We consider the densest submatrix problem, which seeks the submatrix of fixed size of a given binary matrix that contains the most nonzero ent

Real-Time Visual Intelligence on Low-Cost UAVs: A Modular Approach for Tracking, Scanning, and Navigation

AgentsDGX agent

arXiv:2607.02298v1 Announce Type: new Abstract: Autonomous drones are rapidly transforming modern warfare and civil applications alike. This paper presents the development of an integrated intelligent

Token Geometry

Model ReleasesDGX agent

arXiv:2607.01455v1 Announce Type: cross Abstract: Language models learn continuous programs over discrete symbols, with the embedding table and LM-head acting as the read/write interface between them.

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

WARP: Weight-Space Analysis for Recovering Training Data Portfolios

Model ReleasesDGX agent

arXiv:2607.01686v1 Announce Type: new Abstract: Foundation models are routinely released to the public, yet the data recipes used to train them -- such as domain mixture weights that determine how dif

What Types of Human-AI Teams Exist?

TutorialsDGX agent

arXiv:2607.02198v1 Announce Type: cross Abstract: Human-AI teaming has received increasing attention in the literature. However, the range of studies conducted in multiple domains make it difficult to

WorkPods agent memory as wikis. @langchain #openwiki @cognition #deepwiki @karpathy #llmwiki @Factory #autowiki @hwchase17 @BraceSproul

AgentsDGX agent

This post discusses using WorkPods agent memory systems organized as wikis, likely exploring how LLM agents can maintain and access persistent knowledge repositories in wiki format for improved contex

2 Jul 2026

A Unified Benchmark for RCM-Constrained Visual Servoing: Modeling-Controller Interaction and Robustness Analysis in Laparoscopic Robots

Model ReleasesDGX agent

arXiv:2607.00030v1 Announce Type: new Abstract: In robot-assisted laparoscopic minimally invasive surgery (MIS), accurate enforcement of the remote center of motion (RCM) constraint is critical for sa

Accelerating Discrete Diffusion Models with Parallel-In-Time Sampling

HardwareDGX agent

arXiv:2607.00773v1 Announce Type: new Abstract: Discrete diffusion models are widely used for learning and generating discrete distributions. As the generation process is inherently sequential, the ac

AFFMAE: Scalable Vision Pre-Training for High-Resolution Microscopy Segmentation on Desktop Hardware

Model ReleasesDGX agent

arXiv:2602.16249v2 Announce Type: replace Abstract: Self-supervised pretraining has transformed computer vision by enabling data-efficient fine-tuning, yet high-resolution pretraining typically requir

Another fascinating paper on LLM Judges. (bookmark it) It's from Amazon, and they show that if you run panels of LLM judges, averaging their…

Model ReleasesDGX agent

Another fascinating paper on LLM Judges. (bookmark it) It's from Amazon, and they show that if you run panels of LLM judges, averaging their scores is a trap. 'Overall, we establish that robust aggreg

Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering

AgentsDGX agent

arXiv:2607.01087v1 Announce Type: cross Abstract: Generative AI is shifting software engineering from a practice organized around scarce implementation effort toward one organized around abundant, low

Computer vision-based neural networks for radioisotope identification in urban environments

Model ReleasesDGX agent

arXiv:2607.00270v1 Announce Type: cross Abstract: Algorithm development for radioisotope identification in mobile urban search scenarios face significant challenges from non-uniform backgrounds, momen

Dual-Confidence Contrastive Decoding for Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2607.00570v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) increasingly requires models to answer questions from multiple retrieved documents, where only some sources are rel

EPC: A Standardized Protocol for Measuring Evaluator Preference Dynamics in LLM Agent Systems

Model ReleasesDGX agent

arXiv:2607.00297v1 Announce Type: cross Abstract: When LLM agents use evaluator feedback to adapt their behavior in closed loops, evaluator biases propagate through the agent's strategy distribution -

Fraud is Not Just Rarity: A Causal Prototype Attention Approach to Realistic Synthetic Oversampling

ApplicationsDGX agent

arXiv:2507.14706v2 Announce Type: replace-cross Abstract: Detecting fraudulent credit card transactions remains a significant challenge, due to the extreme class imbalance in real-world data and the o

From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape

SafetyDGX agent

arXiv:2606.08625v2 Announce Type: replace Abstract: As Large Language Models (LLMs) advance toward open-ended autonomous agents, the mechanisms used to evaluate and guide their behavior must evolve ac

FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data

Model ReleasesDGX agent

arXiv:2507.10540v3 Announce Type: replace Abstract: The rapid advancement of large language models (LLMs) has created a diverse landscape of models, each excelling at different tasks. This diversity d

Google’s Continued Disruption of Malicious Residential Proxy Networks

SafetyDGX agent

Background Today, in coordination with the FBI, Lumen, and others, Google took action against the NetNut residential proxy network, also known as Popa. This action builds on our disruption of the IPID

How Environment and Urbanization Shape Bird Diversity in Sri Lanka

SafetyDGX agent

arXiv:2607.00582v1 Announce Type: cross Abstract: This study presents a comprehensive analysis of bird diversity across Sri Lanka by integrating spatial, temporal, and environmental data. Bird observa

http://ora.ai is super useful. analyzes the 'agent readiness' of your site, and then gives you a prompt for your coding agents to fix (i'm u…

AgentsDGX agent

ORA.ai is a tool that evaluates website 'agent readiness' by analyzing how well-suited a site is for autonomous agent interaction, then generates prompts that coding agents can use to implement necess

IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video

Model ReleasesDGX agent

arXiv:2603.16432v3 Announce Type: replace Abstract: Unsupervised physical parameter estimation from video lacks a common benchmark: existing methods evaluate on non-overlapping synthetic data, the sol

Large language models replicate and predict human cooperation across experiments in game theory

Model ReleasesDGX agent

arXiv:2511.04500v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as decision-making agents in high-stakes domains and as imitators of human behavior in the so

last day at @aiDotEngineer and i'll be at the Expo's poster area explaining the year's best survey paper on Agent Memory (Hu et al), as we d…

AgentsDGX agent

last day at @aiDotEngineer and i'll be at the Expo's poster area explaining the year's best survey paper on Agent Memory (Hu et al), as we did for @latentspacepod's Paper Club live from the floor with

Last night we hosted the BabyAGI x Physical AI Happy Hour in SF with @yoheinakajima

AgentsDGX agent

Yohei Nakajima hosted a networking event in San Francisco bringing together the BabyAGI and Physical AI communities for casual conversation and relationship-building. The event likely focused on discu

Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning

SafetyDGX agent

arXiv:2505.19614v2 Announce Type: replace-cross Abstract: Multimodal learning has seen remarkable progress, particularly with large-scale pre-training across various modalities. Most current approache

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

Model ReleasesDGX agent

arXiv:2607.00890v1 Announce Type: new Abstract: Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic

RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios

Model ReleasesDGX agent

arXiv:2511.18011v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have demonstrated powerful capabilities in general spatial understanding and reasoning. However, their fine

Robust Operational Space Control with Conformal Disturbance Bounds for Safe Redundant Manipulation

SafetyDGX agent

arXiv:2607.00424v1 Announce Type: new Abstract: Redundant robotic manipulators operating in constrained and human-interactive environments require accurate task-space tracking together with rigorous s

SlowBA: An efficiency backdoor attack towards VLM-based GUI agents

AgentsDGX agent

arXiv:2603.08316v3 Announce Type: replace-cross Abstract: Modern vision-language-model (VLM) based graphical user interface (GUI) agents are expected not only to execute actions accurately but also to

Speech Playground: An Interactive Tool for Speech Analysis and Comparison

SafetyDGX agent

arXiv:2607.00418v1 Announce Type: new Abstract: This paper presents Speech Playground, an interactive speech visualization and comparison tool. While existing tools such as Praat are excellent, it can

Starting in early 2025, Elon Musk and the Trump administration began terminating USAID's programs and firing its staff — with Musk himself b…

ApplicationsDGX agent

Starting in early 2025, Elon Musk and the Trump administration began terminating USAID's programs and firing its staff — with Musk himself boasting about 'feeding it into the woodchipper.' One year ag

Structural Enforcement of Statistical Rigor in AI-Driven Discovery: A Functional Architecture

ApplicationsDGX agent

arXiv:2511.06701v3 Announce Type: replace-cross Abstract: AI-Scientist systems risk manufacturing spurious discoveries through uncontrolled multiple testing. We present a functional architecture that

@swyx with the 🐐 takes on making a talk rejigged my whole talk around the thesis slide for AIE world fair

ToolsDGX agent

@swyx with the 🐐 takes on making a talk rejigged my whole talk around the thesis slide for AIE world fair lots of folks prepping talks next week (congrats!). Some thoughts from RLing on thousands of h

@teortaxesTex AI companies are companies, not labs. It is pompous & silly to call them labs.

IndustryDGX agent

Clem Delangue argues that AI companies should not be referred to as 'labs,' characterizing such terminology as pretentious and inaccurate since these organizations operate as commercial enterprises ra

The Course of News Events: A Comparison of Bottom-Up and Top-Down Approaches for Collecting Text-Based Data about Disasters

TutorialsDGX agent

arXiv:2607.00849v1 Announce Type: new Abstract: News articles are an important source of information on disaster impacts and adaptation. A key methodological challenge in socio-environmental studies i

The MMM Data Model -- A Normative Specification for Knowledge Interoperability in a Decentralisable Knowledge Commons

ApplicationsDGX agent

arXiv:2607.00032v1 Announce Type: new Abstract: Many information systems are built around documents: self-contained units optimised for print production and linear reading. While effective for large-s

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology

Model ReleasesDGX agent

arXiv:2607.01033v1 Announce Type: new Abstract: Model organisms (MOs) - language models trained to exhibit undesired or unnatural behaviours - are frequently used as testbeds for evaluating white-box

This is a very bold letter from Maria Ressa and Bengio! It is historic to see the chairs of the UN Independent International Scientific Pane…

SafetyDGX agent

Maria Ressa and Yann Bengio, prominent figures in technology and AI governance, co-authored a significant letter addressing concerns related to the UN Independent International Scientific Panel on AI.

Understanding How Humans Inject Knowledge into Machine Learning Workflows through Visual Analytics

Model ReleasesDGX agent

arXiv:2607.00969v1 Announce Type: cross Abstract: Visual analytics (VA) plays an increasingly important role in supporting machine learning (ML) workflows. In the field of visualization, such approach

1 Jul 2026

1/ DSGym: A Holistic Framework for Evaluating and Training Data Science Agents Paper: https://arxiv.org/abs/2601.16344

ToolsDGX agent

DSGym is a comprehensive framework designed to evaluate and train AI agents for data science tasks, providing a structured environment for benchmarking agent performance across various data science wo

5/ V1: Unifying Generation and Self-Verification for Parallel Reasoners Paper: https://arxiv.org/abs/2603.04304

ToolsDGX agent

This paper presents V1, a framework that unifies text generation with self-verification mechanisms to enable parallel reasoning processes in language models. The approach allows models to generate mul

A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems

SafetyDGX agent

arXiv:2606.31639v1 Announce Type: cross Abstract: Large language models are no longer only text generators. They are increasingly embedded in retrieval pipelines, enterprise assistants, coding environ

Accelerate protein design with BoltzGen on Amazon SageMaker AI

TutorialsDGX agent

In this post, we demonstrate how to deploy BoltzGen on SageMaker AI and run an end-to-end protein design experiment. By the end of the walkthrough, you have a working setup that scales from quick vali

Agentic-Ideation: Sample Efficient Agentic Trajectories Synthesis for Scientific Ideation Agents

AgentsDGX agent

arXiv:2606.31229v1 Announce Type: new Abstract: Ideation plays a pivotal role in scientific discovery. Recent LLM, especially AI Scientist systems, show promising potential for automated ideation. How

AI systems that are accessible or even better open-source are safer for a simple reason: more people can inspect them, test them, stress the…

SafetyDGX agent

AI systems that are accessible or even better open-source are safer for a simple reason: more people can inspect them, test them, stress them, and report what breaks or harms to keep the builders acco

Building a Multimodal Dataset of Academic Paper for Keyword Extraction

TutorialsDGX agent

arXiv:2606.31069v1 Announce Type: new Abstract: Up to this point, keyword extraction task typically relies solely on textual data. Neglecting visual details and audio features from image and audio mod

Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas?

AgentsDGX agent

arXiv:2606.31213v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed as moral advisors and agents, they need to address dilemmas between two competing values. Ho

CoMNet: A MedNeXt-CorrDiff Framework for Multi-Site Brain Tumor Segmentation

TutorialsDGX agent

arXiv:2606.15305v2 Announce Type: replace Abstract: Accurate brain tumor segmentation from multiparametric magnetic resonance imaging (MRI) is critical for treatment planning, response assessment, and

CSTrader: A Testbed for Language-Grounded Trading in a Community-Driven Virtual Asset Market

Model ReleasesDGX agent

arXiv:2606.31461v1 Announce Type: new Abstract: Niche asset markets, such as Counter-Strike 2 (CS2) weapon skins, are small, volatile, and heavily driven by community discussions and platform rules. T

EgoCogNav: Cognition-aware Human Egocentric Navigation

ApplicationsDGX agent

arXiv:2511.17581v3 Announce Type: replace-cross Abstract: Modeling the cognitive and experiential factors of human navigation is central to deepening our understanding of human-environment interaction

ElemeNet: Multiscale Molecular Machine Learning with Uncertainty Quantification Across the Periodic Table

Model ReleasesDGX agent

arXiv:2606.30961v1 Announce Type: cross Abstract: Advances in deep learning architectures and representations have enabled ML-driven chemical property prediction, but state-of-the-art (SOTA) models ha

Evidence Triangulation for Multimodal Fact-Checking in the Wild

Model ReleasesDGX agent

arXiv:2606.31367v1 Announce Type: cross Abstract: The proliferation of multimedia content on social platforms has fueled multimodal misinformation, where images are used to reinforce false claims. Con

FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge

Model ReleasesDGX agent

arXiv:2602.06625v2 Announce Type: replace Abstract: Existing LLM-as-a-Judge systems suffer from three fundamental limitations: limited adaptivity to task- and domain-specific evaluation criteria, syst

From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching

AgentsDGX agent

arXiv:2601.23088v2 Announce Type: replace-cross Abstract: Semantic caching has emerged as a pivotal technique for scaling LLM applications, widely adopted by major providers including AWS and Microsof

HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

Model ReleasesDGX agent

arXiv:2606.31179v1 Announce Type: new Abstract: As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward

Improved Immiscible Diffusion: Accelerate Diffusion Training by Reducing Its Miscibility

TutorialsDGX agent

arXiv:2505.18521v2 Announce Type: replace Abstract: The substantial training cost of diffusion models hinders their deployment. Immiscible Diffusion recently showed that reducing diffusion trajectory

Large Databases Need Small, Open-Weight Language Models

Local AiDGX agent

arXiv:2606.31808v1 Announce Type: new Abstract: Language model systems built around proprietary APIs often operate on a token-based cost model. This becomes prohibitively expensive in the context of l

Loc2Repair: A Framework for Evaluating the Impact of File-Level Issue Localization in Repo-Level LLM Repair

Local AiDGX agent

arXiv:2606.30963v1 Announce Type: cross Abstract: Repository-grounded automated repair is often reported as a single end-to-end capability, which hides distinct failure modes such as poor file targeti

LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish

Model ReleasesDGX agent

arXiv:2606.31947v1 Announce Type: new Abstract: State-of-the-art speech datasets predominantly focus on widely spoken languages, often overlooking low-resource languages such as Luxembourgish, which r

OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation

Model ReleasesDGX agent

arXiv:2606.31993v1 Announce Type: new Abstract: While robotic manipulation capabilities have advanced rapidly, physical safety remains a major barrier to deploying household robots: task success is in

← Previous
1…384385386387388…428
Next →