AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,499 results
28 Jul 2026

NeurIPS 2026 Reviewer: AI-Generated Rebuttals (and Paper) [D]

Model ReleasesDGX agent

One of the papers I reviewed has what seems to be entirely LLM-generated rebuttals, and the original paper is also clearly LLM-generated, with Claude-speak everywhere. While the authors acknowledge LL

Neuromorphic Object Detection: An In-Depth Study and Future Directions

Model ReleasesDGX agent

arXiv:2607.23576v1 Announce Type: new Abstract: Conventional frame-based cameras face significant challenges in detecting objects under high-speed motion blur or in low-light environments. Neuromorphi

New research from Meta and CMU. This one is on agentic context management for long horizon tasks. (bookmark it) Production agents accumulate…

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

New research from Meta and CMU. This one is on agentic context management for long horizon tasks. (bookmark it) Production agents accumulate context every turn. The usual fix compresses on a token thr

No Optimal Language Set Exists for Multilingual Instruction Tuning: Insights from a Linguistically-Informed Study

Model ReleasesDGX agent

arXiv:2410.07809v2 Announce Type: replace Abstract: Multilingual instruction tuning (MIT) is challenged by the curse of multilinguality, data scarcity, and high computational cost. A natural hypothesi

Not All LLM Reasoning is Visible in the Chain-of-Thought

Model ReleasesDGX agent

arXiv:2607.22925v1 Announce Type: cross Abstract: A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete failure mode w

Nova3D: Code-Native Generation of Programmable 3D Assets

Model ReleasesDGX agent

arXiv:2607.22738v1 Announce Type: cross Abstract: Current 3D generative models mostly produce a final surface: a visually strong but largely opaque mesh. Interactive 3D worlds need more than a surface

Novel Claim or Deja Vu? Rethinking 'Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking

Model ReleasesDGX agent

arXiv:2607.23514v1 Announce Type: cross Abstract: Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks

NSL-SLAM: High-Fidelity Neural Structured-Light Depth for Practical SLAM and Reconstruction

Model ReleasesDGX agent

arXiv:2607.24495v1 Announce Type: new Abstract: Structured-light (SL) cameras power depth sensing in millions of devices, and recent neural SL decoding methods have substantially improved their depth

Numerical Investigation of Sequence Modeling Theory using Controllable Memory Functions

Model ReleasesDGX agent

arXiv:2506.05678v3 Announce Type: replace Abstract: The evolution of sequence modeling architectures, from recurrent neural networks and convolutional models to Transformers and structured state-space

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness

Model ReleasesDGX agent

arXiv:2607.23537v1 Announce Type: new Abstract: Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions

On the Impossibility of Unbiased and Length-Invariant Policy Optimization with Outcome Rewards

Model ReleasesDGX agent

arXiv:2607.23364v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) is the dominant reinforcement learning algorithm for training reasoning capabilities in large language models,

On the Order-Conditional Optimality of Gaffke's Bound

Model ReleasesDGX agent

arXiv:2607.22971v1 Announce Type: cross Abstract: Let X = (X_1, ldots, X_n) be a random vector from any Borel probability law on R_+^n. We revisit the problem of deriving a lower confidence bound (LCB

Online Fair Division with Budget Constraints

Model ReleasesDGX agent

arXiv:2607.23310v1 Announce Type: cross Abstract: We study an online variant of discrete fair division under generalized assignment budget constraints. Goods arrive one at a time and must be assigned

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

Model ReleasesDGX agent

arXiv:2509.02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions

OrchNAS: Orchestrated Neural Architecture Search Service for Personalised Federated Edge Intelligence

Model ReleasesDGX agent

arXiv:2607.22805v1 Announce Type: cross Abstract: We propose OrchNAS, an energy-aware, personalised, federated edge intelligence framework that leverages a Neural Architecture Search Service to automa

OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows

Model ReleasesDGX agent

arXiv:2510.24411v3 Announce Type: replace Abstract: Computer-using agents powered by Vision-Language Models (VLMs) have demonstrated human-like capabilities in operating digital environments like mobi

Out-of-Length Scene Text Recognition: A Two-Axis Diagnosis and a Training-Free Fix

Model ReleasesDGX agent

arXiv:2607.23194v1 Announce Type: new Abstract: Scene Text Recognition (STR) models are trained almost exclusively on word crops of at most 25 characters, yet real deployments (signage, product labels

PAC-DP: PAC-Bayesian Diffusion Policy Learning

Model ReleasesDGX agent

arXiv:2607.24296v1 Announce Type: new Abstract: Diffusion Policies (DPs) are able to perform complex manipulation tasks. However, DPs are typically trained by minimizing a denoising objective, which p

PANOPTICON: A PII-Based Assemblage of Naturalistic Output Tokens for Investigating Privacy Leakage Within LLM Context Window

Model ReleasesDGX agent

arXiv:2607.22695v1 Announce Type: new Abstract: Large Language Models (LLMs) are capable of generalizing human language for the completion of never-before-seen tasks, leading to widespread deployment.

[PAPER] GPQA, MMLU-Pro, and MMMU-Pro were audited for broken questions, and up to 12% of them had to be removed. New drop in clean versions released

Model ReleasesDGX agent

I was very curious why all the models were topping out on GPQA-Diamond around 92 or 93% (AA) and spent the last few weeks pouring over GPQA (Diamond and Extended), and then expanded to auditing MMLU-P

Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation

Model ReleasesDGX agent

arXiv:2607.23694v1 Announce Type: new Abstract: Efficient surgical segmentation empowers clinical diagnosis, intraoperative monitoring, and downstream robotic pipelines for reconstruction and simulati

ParasGB: A Graph Benchmark Suite for Parasitic Estimation on AMS Circuits

Model ReleasesDGX agent

arXiv:2607.23225v1 Announce Type: new Abstract: As chip manufacturing processes advance to deep submicron nodes, parasitic interconnect effects increasingly dominate the performance of analog and mixe

ParBench: A Benchmark for Reliable Evaluation of LLM Parallel Code Translation

Model ReleasesDGX agent

arXiv:2607.22588v1 Announce Type: new Abstract: Modern compute-intensive software must migrate across a changing ecosystem of accelerators, programming APIs, compiler stacks, and portability layers, i

PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis

Model ReleasesDGX agent

arXiv:2607.23794v1 Announce Type: cross Abstract: Pathological diagnosis is inherently multi-scale, requiring the integration of global tissue architecture at low magnification with cellular morpholog

PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs

Model ReleasesDGX agent

arXiv:2607.22859v1 Announce Type: new Abstract: Mathematical Word Problems (MWPs) are an important benchmark for evaluating natural language understanding and quantitative reasoning. Despite recent pr

PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms

Model ReleasesDGX agent

arXiv:2603.27476v2 Announce Type: replace Abstract: AI-powered people search platforms are increasingly used in recruiting, sales prospecting, and professional networking, yet no widely accepted bench

Perplexity brings its Personal Computer AI agent to Windows

Model ReleasesDGX agent

Perplexity AI Inc. today released a Windows version of Personal Computer, expanding its agentic automation software beyond the original Macintosh platform and making it available to more than 1 billio

Phenology-based learning framework for yield estimation and harvest forecasting of raspberry fruits

Model ReleasesDGX agent

arXiv:2411.00967v2 Announce Type: replace Abstract: The future of agriculture is intertwined with automation. Accurate fruit detection, yield estimation, and harvest time prediction are crucial for ef

PhononBench-MP40: a spectrum-resolved benchmark dataset for phonon stability

Model ReleasesDGX agent

arXiv:2607.22573v1 Announce Type: new Abstract: Imaginary phonon modes remain a practical bottleneck in computational materials screening because otherwise plausible structures can be locally dynamica

Physics-Informed Neural Networks for Discovering Periodic Orbits in the Gravitational Three-Body Problem

Model ReleasesDGX agent

arXiv:2607.23501v1 Announce Type: new Abstract: Locating periodic solutions of chaotic dynamical systems normally requires an initial guess close enough to the target orbit for numerical continuation

PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention

Model ReleasesDGX agent

arXiv:2607.24593v1 Announce Type: new Abstract: Token-level sparse attention, as implemented by DeepSeek Sparse Attention (DSA) in production systems, makes the downstream attention efficient but shif

Pointer-Augmented Autoregressive Generation of Patent Claims with Joint Topology and Content Decoding

Model ReleasesDGX agent

arXiv:2607.24040v1 Announce Type: new Abstract: Autoregressive decoders emit flat token sequences and cannot enforce hierarchical constraints across output segments, a limitation that becomes acute in

Poster: Rethinking Security in LLM Code Generation through Real-World Risk Scenarios

Model ReleasesDGX agent

arXiv:2607.23088v1 Announce Type: cross Abstract: Large Language Models (LLMs) are widely used for code generation, yet their security behavior in realistic development workflows remains underexplored

Practical advantage beyond the quadratic speedup limit with fully-quantum walks

Model ReleasesDGX agent

arXiv:2607.22818v1 Announce Type: cross Abstract: We introduce a new class of fully-quantum Metropolis walks in which both the proposal and acceptance steps are intrinsically quantum. Unlike standard

PriSAR: 3D Geometric-Prior-Guided Diffusion for Parameter-Controlled SAR Image Generation

Model ReleasesDGX agent

arXiv:2607.22963v1 Announce Type: cross Abstract: Synthetic aperture radar (SAR) image generation can mitigate data scarcity, but controllablegeneration under sparse observation angles remains difficu

ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents

Model ReleasesDGX agent

arXiv:2606.18037v2 Announce Type: replace Abstract: Tool-using LLM agents increasingly use the Model Context Protocol (MCP) to answer from heterogeneous evidence sources, including search, APIs, datab

proxymate: Diagnosis and Adjustment of Proxy Estimates for Reliable Inference

Model ReleasesDGX agent

arXiv:2607.24401v1 Announce Type: cross Abstract: Proxy outcomes (such as short-term behavioral signals, model predictions, or surrogate endpoints) are frequently used in place of primary outcomes tha

QFoldAgent: An Autonomous Quantum Optimization Multi-Agent System for Protein Structure Prediction

Model ReleasesDGX agent

arXiv:2607.22549v1 Announce Type: new Abstract: Hybrid quantum-classical protein structure prediction depends strongly on Hamiltonian penalty weights, yet existing lattice-based workflows typically fi

Random Forest-Based Prediction of Bone Volume Fraction and Fracture Position from S-Parameters

Model ReleasesDGX agent

arXiv:2607.23563v1 Announce Type: new Abstract: In this paper, we propose a method for predicting bone volume fraction (BVF) and fracture position by constructing a random forest model based on multic

RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning

Model ReleasesDGX agent

arXiv:2607.23290v1 Announce Type: new Abstract: Rare diseases collectively affect an estimated 3.5% to 5.9% of the population, yet more than 70% of patients are misdiagnosed and many endure years of e

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

Model ReleasesDGX agent

arXiv:2607.23927v1 Announce Type: new Abstract: A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, this capacity i

Reason Popper-ly: Patching In-Context Reasoning with Inductive Logic Programming

Model ReleasesDGX agent

arXiv:2607.23019v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting enables large language models (LLMs) to tackle multi-step reasoning tasks, yet the generated intermediate steps are not

Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?

Model ReleasesDGX agent

arXiv:2607.23440v1 Announce Type: cross Abstract: In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists th

Recently I've flipped from being bullish to being bearish about AI. I think I'm updating my bearishness to be more solidly bearish. Early th…

Model ReleasesDGX agent

Recently I've flipped from being bullish to being bearish about AI. I think I'm updating my bearishness to be more solidly bearish. Early thoughts (which I hope to be disproven in the next year or so,

Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models

Model ReleasesDGX agent

arXiv:2601.02580v2 Announce Type: replace-cross Abstract: Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field testing

Reference Feature Atlases for Mechanistic Auditing of Language Models

Model ReleasesDGX agent

arXiv:2607.22570v1 Announce Type: new Abstract: Auditing a new language model usually means relearning and reinterpreting its internal features from scratch. We propose a reference feature atlas: a sp

Rethinking Expert Training for Model Merging with Prompt Learning

Model ReleasesDGX agent

arXiv:2607.24465v1 Announce Type: new Abstract: Model merging aims to combine multiple domain-specialized experts trained from a shared foundation model into a single multi-task model. Existing approa

Risk Governance for Generative AI Mental Health Support: A Multi-Turn Safety Architecture

Model ReleasesDGX agent

arXiv:2607.22692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existi

Risk reversal for least squares estimators under nested convex constraints

Model ReleasesDGX agent

arXiv:2601.16041v2 Announce Type: replace-cross Abstract: In constrained stochastic optimization, one expects that restricting the feasible set, provided it still contains the true parameter, should n

Robustifying pathology foundation models via fine-tuning

Model ReleasesDGX agent

arXiv:2607.22861v1 Announce Type: cross Abstract: Pathology foundation models (FMs) produce powerful tile-level representations which remain sensitive to scanner and staining variability, undermining

RRTrack: Robust and Recoverable Object 6D Pose Tracking for Dynamic Scenes

Model ReleasesDGX agent

arXiv:2607.23669v1 Announce Type: new Abstract: Robust object 6D pose tracking is critical for robotic systems operating in dynamic and occluded scenes. Per-frame estimators are accurate but computati

SAFE-MEME: Structured Reasoning Framework for Robust Hate Speech Detection in Memes

Model ReleasesDGX agent

arXiv:2412.20541v2 Announce Type: replace Abstract: Memes act as cryptic tools for sharing sensitive ideas, often requiring contextual knowledge to interpret them correctly. It makes multimodal meme m

SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems

Model ReleasesDGX agent

arXiv:2603.03536v2 Announce Type: replace-cross Abstract: Current LLM-based conversational recommender systems (CRS) primarily optimize recommendation accuracy and user satisfaction. We identify an un

SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI

Model ReleasesDGX agent

arXiv:2607.22926v1 Announce Type: new Abstract: High-impact generative AI makes catastrophic misuse a lifecycle-control problem, not merely a prompt-filtering problem. SAGE is a safety-first, authoriz

SCAIR: Schema-Conditioned Agentic Iterative Reasoning for Enterprise Knowledge Graphs

Model ReleasesDGX agent

arXiv:2607.22571v1 Announce Type: new Abstract: Knowledge Graph-based Retrieval-Augmented Generation (KG-RAG) enables natural language interaction with structured enterprise knowledge, yet existing ag

Scale Weight Decay and Train Better

Model ReleasesDGX agent

arXiv:2607.23777v1 Announce Type: cross Abstract: The discovery of scaling laws has motivated training neural networks on ever increasing quantities of data. This is typically done with a constant dec

scMIR: a vision-language foundation model for single-cell light microscopy image representation

Model ReleasesDGX agent

arXiv:2607.22712v1 Announce Type: cross Abstract: Single-cell light microscopy images have become an important data source for characterizing cell phenotypes, but their complexity and heterogeneity po

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling

Model ReleasesDGX agent

arXiv:2510.14717v2 Announce Type: replace-cross Abstract: Increasing the batch size during training -- a ''batch ramp'' -- is a promising strategy to accelerate large language model pretraining. While

SEGRA: Structured Experience-Guided Graph Reasoning Agent for Gremlin Based Question Answering

Model ReleasesDGX agent

arXiv:2607.22713v1 Announce Type: new Abstract: Enterprise IT support knowledge graphs capture rich relationships among cases, users, devices, symptoms, taxonomic categories, root causes, and historic

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B

Model ReleasesDGX agent

arXiv:2607.22545v1 Announce Type: cross Abstract: Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, re

← Previous
1…6263646566…375
Next →