AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,428
  • Agents7,398
  • Applications5,301
  • Concepts5
  • Hardware1,785
  • Industry6,113
  • Local Ai4,833
  • Model Releases23,177
  • Research19,713
  • Safety13,092
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,428
  • Agents7,398
  • Applications5,301
  • Concepts5
  • Hardware1,785
  • Industry6,113
  • Local Ai4,833
  • Model Releases23,177
  • Research19,713
  • Safety13,092
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

86,428Total entries
1Added by human
86,427Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,016 results
4 Aug 2026

LLM-OSDA: An Optimal-Stopping Dynamic Auction for Native Advertising in Multi-Turn LLM Conversations

SafetyDGX agent

arXiv:2608.00123v1 Announce Type: new Abstract: LLM-native advertising embeds sponsored content directly into model-generated responses, shifting the unit of sale from a fixed slot to a moment within

LUT: Latent Utility Training for Visual Reasoning

SafetyDGX agent

arXiv:2608.00743v1 Announce Type: new Abstract: Multimodal large language models have advanced visual understanding, yet perception-intensive reasoning remains challenging. Recent latent visual reason

MemoryForge: Synthesize Lifelong Memory for Human-Like LLM Agents

AgentsDGX agent

arXiv:2608.00007v1 Announce Type: new Abstract: Equipping Large Language Models (LLMs) with human-like personas is crucial for agentic applications, such as role-play and user simulation. Traditional

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Messages, Not Tokens: Grounded Coresets for Faithful VLM Compression

ResearchDGX agent

arXiv:2608.02134v1 Announce Type: new Abstract: Modern vision language models (VLMs) turn high-resolution images into long sequences of visual tokens. Every token traverses the language decoder and pe

MonitorVLM-v2: A Deployed Vision-Language Framework for Real-Time Safety Violation Detection

SafetyDGX agent

arXiv:2608.00975v1 Announce Type: new Abstract: Large vision--language models (VLMs) can reason step by step about complex visual scenes, but this open-ended, autoregressive chain-of-thought (CoT) app

Multi-View Unified Camera Fields: Geometry-Shaped Action-Facing Representations for RGB-Only Multi-Camera VLA Policies

ApplicationsDGX agent

arXiv:2608.01826v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation, yet complex contact-rich tasks often benefit from multi-ca

Neural Circuit Function Inference with LLMs

ResearchDGX agent

arXiv:2608.00059v1 Announce Type: new Abstract: The success of connectome mapping now shifts the challenge of understanding the nervous system to the interpretation of neural circuits. Here, we devise

Nonparametric Distribution Regression Re-calibration

SafetyDGX agent

arXiv:2602.13362v2 Announce Type: replace-cross Abstract: A key challenge in probabilistic regression is ensuring that predictive distributions accurately reflect true empirical uncertainty. Minimizin

One-Sided Quantile Coupling for Flow Matching

SafetyDGX agent

arXiv:2608.00978v1 Announce Type: cross Abstract: Flow Matching trains continuous-time generative models by regressing the velocity field of a probability path between a simple source distribution and

Partial FC: Training 10 Million Identities on a Single Machine

HardwareDGX agent

arXiv:2010.05222v3 Announce Type: replace Abstract: Training face recognition models with millions of identities is challenging because classifier storage, logit memory, and computation grow linearly

PhenoStitch: Training-Free Panoptic Crop Mapping from Satellite Image Time Series

ResearchDGX agent

arXiv:2608.00870v1 Announce Type: new Abstract: Panoptic crop mapping requires both delineating individual agricultural parcels and assigning a crop type to each parcel from satellite image time serie

PixelSR: Efficient Screen Content Super-Resolution via Pixel Classification

Local AiDGX agent

arXiv:2608.00646v1 Announce Type: new Abstract: Screen content images are generally composed of texts and graphics. Compared to natural images, these man-made images contain a large quantity of sharp

QuASH: Using Natural-Language Heuristics to Query Visual-Language Robotic Maps

ResearchDGX agent

arXiv:2510.14546v2 Announce Type: replace Abstract: Embeddings from Visual-Language Models are increasingly utilized to represent semantics in robotic maps, offering an open-vocabulary scene understan

RadYOLO: Computationally Efficient 3D Object Detection and Segmentation in CT and MRI

Local AiDGX agent

arXiv:2608.00508v1 Announce Type: new Abstract: Object detection and segmentation in three-dimensional medical images is a very active area of research. However, most proposed deep learning models car

Rolling Shutter Camera Self-Calibration

ResearchDGX agent

arXiv:2608.01509v1 Announce Type: new Abstract: Rolling shutter (RS) cameras are widely used in consumer devices, but their row-wise exposure causes distortions under motion, making geometric 3D visio

SeDeM: Selective Decompression of Hidden-State Memories for Long-Context Question Answering

ResearchDGX agent

arXiv:2608.00311v1 Announce Type: new Abstract: Long-context inference with large language models (LLMs) is costly: self-attention during prefill scales quadratically with sequence length, and the key

Select-And-Extract: A Lightweight Plugin for Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2608.00658v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) for language model (LM) systems fundamentally has two failure modes: retrieval failure and reading failure. The for

SG-Layout: Structured Scene Graph-Guided Layout Generation with LLMs

SafetyDGX agent

arXiv:2608.01106v1 Announce Type: new Abstract: Understanding and generating spatially coherent layouts from natural language remains a fundamental yet challenging task for large language models (LLMs

SIEVE: Selective Integrity Verification and Escalation for Defending LLM Agents against Indirect Prompt Injection

AgentsDGX agent

arXiv:2512.06716v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used as the core of agentic systems due to their strong reasoning, planning, and tool-use capabi

Start Classifying: Categorical Critics for LLM Reinforcement Learning

SafetyDGX agent

arXiv:2608.02181v1 Announce Type: new Abstract: Proximal Policy Optimization (PPO) for large language models typically trains its critic by mean-squared-error (MSE) regression on scalar value targets.

Towards Compact Unified Multimodal Tracking: Synergizing Knowledge Distillation with Structural Pruning

SafetyDGX agent

arXiv:2608.01488v1 Announce Type: new Abstract: Unified multimodal object tracking has achieved remarkable robustness by leveraging complementary sensor data (e.g., RGB, Thermal, Depth), yet the heavy

Towards General Language-Conditioned Latent Safety Filters

SafetyDGX agent

arXiv:2608.00315v1 Announce Type: cross Abstract: Robot policies are becoming increasingly general, with vision-language-action (VLA) models enabling a single policy to execute diverse tasks specified

UAV-Based Environmental Monitoring of Rip-Current Indicators Using Wavelet-Derived Texture Features

SafetyDGX agent

arXiv:2608.02448v1 Announce Type: new Abstract: Rip currents are recurrent coastal natural hazards that threaten beachgoers and create operational challenges for lifeguards and coastal managers. Relia

UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization

ResearchDGX agent

arXiv:2608.01706v1 Announce Type: new Abstract: Recent geometric foundation models enable feed-forward inference for SLAM, but their predictions are strongly dependent on the input view set, which lea

Unpacking ChatGPT Work: the Agent for a Billion Users

AgentsDGX agent

ChatGPT Work was launched by OpenAI on July 9, 2026 as an agent‑oriented knowledge‑work platform that combines chat, Codex tools and cloud agents across fourteen model configurations. Within three wee

Using Non-Lipschitz Signum-based Functions for Distributed Optimization and Machine Learning: Trade-off Between Con-vergence Rate and Optimality Gap

ResearchDGX agent

arXiv:2608.01220v1 Announce Type: cross Abstract: In recent years, the prevalence of large-scale data-sets and the demand for sophisti-cated learning models have necessitated the development of effici

VertiAKD: Adaptive Off-Road Kinodynamics on Vertically Challenging Terrain

Local AiDGX agent

arXiv:2608.00945v1 Announce Type: new Abstract: Off-road mobility requires autonomous mobile robots to generalize across heterogeneous vehicle fleets and continuously changing terrain conditions. Exis

VGER: Voxel-Guided Global Event Ranking for Event Cloud Attribution

ResearchDGX agent

arXiv:2608.01470v1 Announce Type: new Abstract: Event cameras produce sparse and asynchronous event streams that provide rich spatio-temporal information for efficient perception. Recent advances in e

When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems

AgentsDGX agent

arXiv:2608.00747v1 Announce Type: new Abstract: Large language models are increasingly integrated into autonomous robotic systems for task planning and control, but this integration exposes them to pr

Where Does Generative Difficulty Reside? An Empirical Study of Target Representations

Local AiDGX agent

arXiv:2608.00626v1 Announce Type: new Abstract: The target representation defines the distribution an image generator must learn, yet it is often treated as an interchangeable interface. This assumpti

WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting

ResearchDGX agent

arXiv:2510.10726v2 Announce Type: replace Abstract: We present WorldMirror, a unified feed-forward model for comprehensive 3D geometric prediction tasks. Unlike existing methods constrained to image-o

3 Aug 2026

ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency

SafetyDGX agent

arXiv:2607.29169v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies achieve strong performance in robotic manipulation but remain vulnerable to runtime disturbances that break the

Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning

SafetyDGX agent

arXiv:2607.28986v1 Announce Type: cross Abstract: Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only corpora and

AgenticRepair: Multi-Faceted Program Context Engineering for Agentic Vulnerability Repair

AgentsDGX agent

arXiv:2607.29422v1 Announce Type: cross Abstract: Automated vulnerability repair aims to reduce the time and effort required to patch security flaws from a vulnerability triage report. Recent agentic

Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth

SafetyDGX agent

arXiv:2607.28890v1 Announce Type: cross Abstract: Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that presumes h

Artifact detection and localization in single-channel mobile EEG for sleep research using deep learning and attention mechanisms

Local AiDGX agent

arXiv:2504.08469v3 Announce Type: replace-cross Abstract: Current methods for detecting artifacts in sleep EEG range from threshold-based algorithms to machine learning approaches, yet applications re

Behind the scenes: How we build, test, and scale Google Agent Skills

AgentsDGX agent

AI agents are only as good as the instructions and context you give them. When we launched Google Agent Skills, our goal was simple: encode Google Cloud domain knowledge into structured, open-source i

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

Local AiDGX agent

arXiv:2607.28818v1 Announce Type: new Abstract: As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do

Bridging the Question-Answer Gap in Retrieval-Augmented Generation: Hypothetical Prompt Embeddings

SafetyDGX agent

arXiv:2607.29402v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems synergize retrieval mechanisms with generative language models to enhance the accuracy and relevance of r

Conditioning Tree-Based Diffusions and Flows for Probabilistic Tabular Regression

ResearchDGX agent

arXiv:2607.28864v1 Announce Type: cross Abstract: Tree-based diffusion models fit flexible conditional predictive distributions for tabular regression without a neural density estimator, but they inhe

Cross-Lingual Transfer for Machine Translation in Turkic Languages

ResearchDGX agent

arXiv:2607.29355v1 Announce Type: cross Abstract: Cross-lingual transfer is central to low-resource machine translation, but its behavior within closely related language families remains insufficientl

Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning

Local AiDGX agent

arXiv:2607.22186v2 Announce Type: replace Abstract: Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimiza

Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation

SafetyDGX agent

arXiv:2603.13415v2 Announce Type: replace Abstract: Valence-arousal (VA) estimation is crucial for capturing the nuanced nature of human emotions in naturalistic environments. While pre-trained vision

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL

SafetyDGX agent

arXiv:2607.29246v1 Announce Type: new Abstract: Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a

DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs

SafetyDGX agent

arXiv:2607.29112v1 Announce Type: cross Abstract: Audio-visual speech recognition (AVSR) relies on effective fusion of audio and visual modalities, yet existing approaches treat cross-modal interactio

DuetHOI: Language-Guided Bimanual Hand--Object Motion Generation with Articulation Planning and Contact Refinement

SafetyDGX agent

arXiv:2603.08390v3 Announce Type: replace-cross Abstract: Bimanual articulated-object interaction generation requires a model to capture the evolution of object articulation, coordination between the

Enabling Low-Latency Machine learning on Radiation-Hard FPGAs with hls4ml

ResearchDGX agent

arXiv:2602.15751v2 Announce Type: replace-cross Abstract: This paper presents an end-to-end demonstration of a viable, ultra-fast, radiation-hard machine learning (ML) application on FPGAs, which coul

End-to-End Fairness Optimization with Fair Decision-Focused Learning

SafetyDGX agent

arXiv:2607.29441v1 Announce Type: new Abstract: Many real-world systems rely on predictive models to inform decisions, and fairness concerns arise in both the prediction and decision stages. We introd

Here’s why AI agents lie and cheat to reach their goals

ResearchDGX agent

MIT Technology Review Explains: Let our writers untangle the complex, messy world of technology to help you understand what’s coming next. You can read more from the series here. When two OpenAI model

Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use

AgentsDGX agent

arXiv:2607.28889v1 Announce Type: cross Abstract: Qualitative researchers increasingly encounter interaction corpora whose scale exceeds what manual coding alone can address, and large language models

'I ran my own benchmarks on it' seems to be pretty common comment around here. How about dedicating a thread for this and sharing?

TutorialsDGX agent

Of course, the concern is that in the end, this thread will be fed into the models' training data, but I feel benchmarking isn't so open and very fragmented. submitted by /u/jinnyjuice [link] [comment

Kinodynamic Motion Retargeting for Humanoid Locomotion via Multi-Contact Whole-Body Trajectory Optimization

SafetyDGX agent

arXiv:2603.09956v2 Announce Type: replace Abstract: We present the KinoDynamic Motion Retargeting (KDMR) framework, a novel approach for humanoid locomotion that models the retargeting process as a mu

Knowledge Restoration-driven Prompt Optimization: Unlocking LLM Potential for Open-Domain Relational Triplet Extraction

ResearchDGX agent

arXiv:2601.15037v2 Announce Type: replace-cross Abstract: Open-domain Relational Triplet Extraction (ORTE) aims to mine structured knowledge without predefined relation schemas. Large Language Models

Lightweight Neural Networks for Affordance Segmentation: Enhancement of the Decoder Module

ApplicationsDGX agent

arXiv:2607.29473v1 Announce Type: new Abstract: The deployment of deep neural networks for visual affordance segmentation on wearable robots poses may prove critical, due to some conflicting aspects o

Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory

ResearchDGX agent

arXiv:2607.29167v1 Announce Type: cross Abstract: Long-term memory lets large language model(LLM) agents reuse prior preferences and work flows, but it also turns untrusted observations into persisten

Mirror Learning

SafetyDGX agent

arXiv:2607.28737v1 Announce Type: cross Abstract: We investigate imitation learning through the lens of third-person observation and propose a framework for mirror learning: acquiring actionable polic

MolGVR: A Chemistry-Grounded Framework for Text-to-Molecule Generation

ResearchDGX agent

arXiv:2607.29479v1 Announce Type: new Abstract: Text-to-molecule generation is typically formulated as a one-shot sequence generation problem, where a model directly maps target descriptions to molecu

MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation

TutorialsDGX agent

arXiv:2607.29180v1 Announce Type: cross Abstract: Text-to-motion generation must produce motions that are semantically correct, temporally coherent, and physically plausible. A natural approach is to

NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability

AgentsDGX agent

arXiv:2607.28942v1 Announce Type: new Abstract: Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retrieval-augmented

OsteoCAD: A Human-in-the-Loop Cloud-Edge Framework for Bone Tumor Segmentation

HardwareDGX agent

arXiv:2607.29266v1 Announce Type: cross Abstract: Artificial Intelligence (AI) and Deep Learning (DL) have notably advanced medical image analysis, yet many health- care organizations struggle to adop

← Previous
1…737738739740741…1034
Next →