AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

86,993Total entries
1Added by human
86,992Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,477 results
31 Jul 2026

OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation

Model ReleasesDGX agent

arXiv:2607.27278v1 Announce Type: new Abstract: Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchm

Property-driven Causal Abstractions for Markov Decision Processes

ResearchDGX agent

arXiv:2607.26787v1 Announce Type: new Abstract: Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state variables and th

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.28227v1 Announce Type: cross Abstract: GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision a

RoboBRIDGE: A Modular Framework for Bridging Policies to Robust Real-World Robotic Agents

AgentsDGX agent

arXiv:2607.27881v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have attracted growing interest as a scalable approach to robotic manipulation. While these models are effective act

SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation

Model ReleasesDGX agent

arXiv:2607.26313v1 Announce Type: cross Abstract: Agentic systems act, so a defect in the evidence they retrieve becomes a wrong action with a currency cost. The most dangerous enterprise defects are

Scaling medical imaging report generation with multimodal reinforcement learning

Model ReleasesDGX agent

arXiv:2601.17151v2 Announce Type: replace-cross Abstract: Frontier models have demonstrated remarkable capabilities in understanding and reasoning with natural-language text, but they still exhibit ma

SemAnCorr: Semantic Anchored Correspondence for Zero-Shot Manipulation Skill Transfer

Model ReleasesDGX agent

arXiv:2607.28382v1 Announce Type: new Abstract: Transferring manipulation skills across object instances that share functionality but differ in geometry remains a fundamental challenge in robot learni

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

Model ReleasesDGX agent

arXiv:2607.26314v1 Announce Type: cross Abstract: Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisti

When Should AI Follow? Task Structure and Joint Adaptation by Human and AI Agents

Local AiDGX agent

arXiv:2504.20903v4 Announce Type: replace-cross Abstract: How should organizations divide and sequence decision tasks between human and artificial agents? We develop a computational model of joint seq

Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees

Model ReleasesDGX agent

arXiv:2607.28399v1 Announce Type: new Abstract: Computer-use agents often fail on transient GUI events because they produce the correct action only after the relevant window has already closed. We ide

30 Jul 2026

Advancing the price-performance frontier with GPT‑5.6

Model ReleasesDGX agent

Advancing the price-performance frontier with GPT‑5.6 Huge price drop from OpenAI today: GPT-5.6 Terra got a 20% reduction, and GPT-5.6 Luna got a massive 80% drop. OpenAI credit 5.6 Sol with enabling

Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach

Model ReleasesDGX agent

arXiv:2607.26317v1 Announce Type: cross Abstract: Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a

Batten Down Your Packages: Mitigation Guidance for Supply Chain Compromise

Model ReleasesDGX agent

Written by: Kelli Vanderlee, Stuart Carrera For years, the cybersecurity industry's understanding of software supply chain compromise has been anchored by a few watershed events, including Russian cyb

ContactFlow: A video action conditioning that transfers across embodiments

ApplicationsDGX agent

arXiv:2607.26579v1 Announce Type: cross Abstract: World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. Howe

Crossing-Free Probabilistic K-Line Forecasts Without Retraining

Model ReleasesDGX agent

arXiv:2607.26792v1 Announce Type: cross Abstract: Probabilistic K-line forecasting describes uncertainty in four complementary prices, namely open--high--low--close (OHLC). However, it introduces two

Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text

Model ReleasesDGX agent

arXiv:2607.26368v1 Announce Type: new Abstract: Financial disclosures contain numerical claims, temporal statements, entity references, policy commitments, and risk descriptions that may conflict in q

EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

Model ReleasesDGX agent

arXiv:2607.26518v1 Announce Type: new Abstract: Reliable visual safety understanding in real-world scenarios demands more than just object recognition; it requires causal reasoning under epistemic unc

Enhancing Automated Machine Learning via Homogeneous Train-Test Splitting Methods

Model ReleasesDGX agent

arXiv:2607.26625v1 Announce Type: new Abstract: Accurate model evaluation in machine learning depends critically on how datasets are split into training and testing subsets. Standard random splitting

FedWeave: Rethinking the Unit of Specialization in Heterogeneous Federated MoE-LoRA

Model ReleasesDGX agent

arXiv:2607.26618v1 Announce Type: cross Abstract: Federated PEFT enables LLMs to collaboratively adapt to decentralized private data without sharing raw examples. However, task heterogeneity across cl

Financial Volatility and Risk Forecasting Incorporating a Larger Number of Realized Measures

Model ReleasesDGX agent

arXiv:2411.17136v2 Announce Type: replace-cross Abstract: Realised volatility has become increasingly prominent in volatility forecasting due to its ability to capture intraday price fluctuations. Wit

Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement

ApplicationsDGX agent

arXiv:2607.26473v1 Announce Type: cross Abstract: Personalizing large language models (LLMs) to individual users is essential for improving user experience, yet existing approaches typically rely on e

LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving

Model ReleasesDGX agent

arXiv:2607.26491v1 Announce Type: cross Abstract: The energy consumption of Large Language Model (LLM) serving is becoming a major system challenge as deployment scales, driven by hardware power and t

Mitigating Compounding Error via Video Representation Regularization

AgentsDGX agent

arXiv:2607.27036v1 Announce Type: new Abstract: Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window

Prior Directions: Why GUI Grounding Gets Locked in the Past

ResearchDGX agent

arXiv:2607.26913v1 Announce Type: new Abstract: Vision-language models often use descriptions of earlier visual states to make decisions about the current scene. When the scene changes, stale language

Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method

Model ReleasesDGX agent

arXiv:2607.26924v1 Announce Type: new Abstract: Recent work on LeWorldModel (LeWM) has shown that the Sketched Isotropic Gaussian Regularizer (SIGReg) enables stable end-to-end world-model learning fr

Transformers Can Learn Rules They've Never Seen: Proof of Computation Beyond Interpolation

Model ReleasesDGX agent

arXiv:2603.17019v2 Announce Type: replace Abstract: A central question in the debate over large language models is whether transformers can learn rules they have never seen, or whether they can only i

29 Jul 2026

Addressable Recall Compaction for Long Context-Window Control in AI Agents

Model ReleasesDGX agent

arXiv:2607.25066v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate reasoning traces, actions, and tool observations that can eventually exceed a model's fixed context window. Existing

Agentic AI for Scientific Reasoning in Autonomous Quantum Sensing Experiments

Model ReleasesDGX agent

arXiv:2607.25145v1 Announce Type: cross Abstract: We implement an agentic AI workflow built around a large language model (LLM) agent for autonomous experiments with nitrogen-vacancy (NV) centers in d

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation

Model ReleasesDGX agent

arXiv:2607.25881v1 Announce Type: new Abstract: We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were indepen

At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference

Model ReleasesDGX agent

arXiv:2607.25504v1 Announce Type: cross Abstract: Fine-grained weight pruning and activation sparsification have emerged as effective approaches for reducing the compute and memory cost of inference f

Authoring Agent Skills: A Software-Engineering Approach

Model ReleasesDGX agent

arXiv:2607.25032v1 Announce Type: cross Abstract: Agent Skills are an emerging way to extend large language model agents with reusable procedural knowledge that the agent loads on demand. Anthropic in

A.X-K2 released

Model ReleasesDGX agent

https://huggingface.co/skt/A.X-K2 https://huggingface.co/skt/A.X-K2-ALM https://huggingface.co/KRAFTON/A.X-K2-Raon-Speech-21B-A3B 688B-A33B + About South Korea's Soverign AI Foundation Model Project.

Bridging Compute- and Data-Optimal Pretraining

Model ReleasesDGX agent

arXiv:2607.25271v1 Announce Type: cross Abstract: Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in whic

CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs

Model ReleasesDGX agent

arXiv:2607.24754v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to provide mental health support, requiring reliable evaluation of safety, empathy, and therapeutic

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

Model ReleasesDGX agent

arXiv:2607.25294v1 Announce Type: cross Abstract: Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has hig

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

SafetyDGX agent

arXiv:2607.25659v1 Announce Type: new Abstract: Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines,

COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution

Model ReleasesDGX agent

arXiv:2607.25400v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly entrusted with natural-language workflow instructions (e.g., retail-payment policies) that specify no

Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA

Model ReleasesDGX agent

arXiv:2607.25921v1 Announce Type: cross Abstract: In this work, we study the use of Vision-Language Models (VLMs) for anomaly detection in an agent-driven game Quality Assurance (QA) pipeline focusing

From Training to Deployment: Post-Hoc Causal Feature Identification via Sensitivity Ratios

ResearchDGX agent

arXiv:2607.25546v1 Announce Type: new Abstract: Given a model that is already trained, which features does it rely on causally versus spuriously? Existing methods require access to the training proced

Generative Distributionally Robust Optimization

ResearchDGX agent

arXiv:2607.24983v1 Announce Type: cross Abstract: Generative models are increasingly adopted in distributionally robust optimization (DRO), but existing approaches trade off model compatibility and ad

I pre-trained a 700m on 18B tokens optimized for Python and Wikitext | TheOneWhoWill/Shibai-700M-Base · Hugging Face

Model ReleasesDGX agent

I know this is the 1000000th new sub billion parameter model out there and probably isn't as good as Qwen 3 0.6B or Qwen 3.5 0.8B but it still packs a decent punch. My intention to to continuously pre

IMPRINT: Image-Conditioned Query Enrichment for Long-Tail Object Goal Navigation

Model ReleasesDGX agent

arXiv:2607.25106v1 Announce Type: new Abstract: Embodied AI increasingly relies on queryable semantic maps built from pre-trained vision-language models to enable zero-shot Object Goal Navigation (Obj

Med-R^3: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning

Model ReleasesDGX agent

arXiv:2507.23541v5 Announce Type: replace Abstract: In medical scenarios, effectively retrieving external knowledge and leveraging it for rigorous logical reasoning is of significant importance. Despi

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing

Model ReleasesDGX agent

arXiv:2607.25300v1 Announce Type: new Abstract: Video editing is fundamentally message-driven: even from the same source footage, the selected shots change depending on the narrative the editor wishes

OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2607.25641v1 Announce Type: cross Abstract: While text-to-image models exhibit remarkable visual fidelity, they frequently violate fundamental physical commonsense. Existing benchmarks often rel

Reinforcement Learning for Code Optimization

Model ReleasesDGX agent

arXiv:2607.25970v1 Announce Type: cross Abstract: RL for code correctness is now established: have the model generate a program, run it against hidden test cases, and reward solutions that pass. Exten

Rethinking CD: A Reproducibility Study and Extension on the Ineffectiveness of Contrastive Decoding at Mitigating Object Hallucinations in MLLMs

Model ReleasesDGX agent

arXiv:2607.25196v1 Announce Type: new Abstract: Contrastive decoding (CD) has been proposed as a training-free strategy for mitigating object hallucinations in multimodal large language models (MLLMs)

ScaleResfusion: Residual Rectified Flow based on Residual Vector Field

Model ReleasesDGX agent

arXiv:2607.25275v1 Announce Type: cross Abstract: Real-world Image Restoration (Real-IR) aims to recover high-quality (HQ) images from complex and unknown degradations. Although recent diffusion-based

Two of the people most responsible for scaling the transformer are now betting on a next act. @MillionInt ran the Reasoning 🍓 team at OpenA…

Model ReleasesDGX agent

Two of the people most responsible for scaling the transformer are now betting on a next act. @MillionInt ran the Reasoning 🍓 team at OpenAI. @_arohan_ was a pre-training lead on Gemini after years at

Understand Kimi K3 from first principles: a recommended order for anyone trying to understand this beast

Local AiDGX agent

Everyone is talking about Kimi K3, but if you jump straight into the technical report, you’ll quickly realize it’s standing on years of research -- just like any breakthrough is! If you want to unders

28 Jul 2026

Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Coding Agent Teams

Model ReleasesDGX agent

arXiv:2607.22917v1 Announce Type: new Abstract: Large Language Model (LLM) agents have significantly improved coding and programming workflows. Claude Code, in particular, is one of the most powerful

Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization

ResearchDGX agent

arXiv:2607.24354v1 Announce Type: new Abstract: Automatic prompt optimization (APO) has been widely adopted to adapt vision-language models (VLMs) to downstream tasks without weight updates, yielding

AutoMat: Enabling Automated Crystal Structure Reconstruction from Microscopy via Agentic Tool Use

Model ReleasesDGX agent

arXiv:2505.12650v2 Announce Type: replace-cross Abstract: Reconstructing atomistic crystal structures from a single noisy STEM projection is an ill-posed inverse problem: multiple lattices can explain

Biggest ever MCP update brings metadata, cybersecurity enhancements

Model ReleasesDGX agent

The developers of the Model Context Protocol, an open-source technology that underpins many artificial intelligence applications, today released a new version of the software. The release is described

Breaking the Synthetic-Real Domain Shortcut for Training-Free Generative Replay-based Class Incremental Learning

SafetyDGX agent

arXiv:2607.22994v1 Announce Type: new Abstract: Class-incremental learning (CIL) requires models to continuously acquire new knowledge while avoiding catastrophic forgetting. While exemplar replay is

CameraAnything: Refilming Videos with Arbitrary Camera Control

Model ReleasesDGX agent

arXiv:2607.24591v1 Announce Type: new Abstract: We introduce CameraAnything, the first unified framework for camera controlled video editing that enables joint control of both intrinsic and extrinsic

Decentralized Granular Access Control for Agentic AI Systems in Critical Infrastructure

Model ReleasesDGX agent

arXiv:2607.22611v1 Announce Type: new Abstract: The deployment of autonomous AI agents in production infrastructure introduces fundamental security challenges that traditional role-based access contro

DeepLook: Deeper Thinking with Lookahead

Model ReleasesDGX agent

arXiv:2607.22602v1 Announce Type: new Abstract: Inference-time scaling has emerged as a powerful paradigm for improving large language model reasoning, often delivering larger gains on difficult reaso

Diffusion-Guided Search via Exponential Tilting (DiffTilt): An Application to Falsification of Safety-Critical Systems

Model ReleasesDGX agent

arXiv:2607.23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation. Existing f

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference

Model ReleasesDGX agent

arXiv:2607.24434v1 Announce Type: cross Abstract: Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, bu

← Previous
1…331332333334335…1042
Next →