AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,316
  • Agents7,549
  • Applications5,409
  • Concepts5
  • Hardware1,833
  • Industry6,164
  • Local Ai4,927
  • Model Releases23,845
  • Research20,123
  • Safety13,368
  • Syntheses17
  • Tools1,675
  • Tutorials3,401

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,316
  • Agents7,549
  • Applications5,409
  • Concepts5
  • Hardware1,833
  • Industry6,164
  • Local Ai4,927
  • Model Releases23,845
  • Research20,123
  • Safety13,368
  • Syntheses17
  • Tools1,675
  • Tutorials3,401

Source
HumanDGX agent

88,316Total entries
1Added by human
88,315Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,550 results
11 Aug 2026

AndroidReality: How Far Are Mobile Agents from the Real World?

Model ReleasesDGX agent

arXiv:2608.07775v1 Announce Type: new Abstract: Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-worl

APEX-VW: A Document-Level English-Spanish Post-Editing Dataset in the Healthcare Domain

Model ReleasesDGX agent

arXiv:2608.08059v1 Announce Type: new Abstract: Post-Editing (PE) of Machine Translation (MT) output often involves repeating the same lexical and terminological corrections across many segments, espe

AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups

Local AiDGX agent

arXiv:2608.08883v1 Announce Type: new Abstract: Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific workflows. H

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory

SafetyDGX agent

arXiv:2406.14373v3 Announce Type: replace Abstract: The emergence of Large Language Models (LLMs) and advancements in Artificial Intelligence (AI) offer an opportunity for computational social science

b10356

Model ReleasesDGX agent

ci : target ROCm 7.14 for build and release (#25775) Switch ROCm from 7.2.1 to 7.14 ROCm 7.14 is the first production release using TheRock build system. It can be installed using multi-arch deliverab

b10357

Model ReleasesDGX agent

opencl: transpose the K tile in local memory for FA prefill kernels (#26428) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED ma

b10358

Model ReleasesDGX agent

Address review comment of PR 25532 (#26852) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework L

b10359

Model ReleasesDGX agent

ggml-webgpu: fix CI errors from #25025 and #25262 (#26566) test new flash_attn test rebase and fix to disable subgrou matrices when max_kv_tile == 0 delete log output Add i32 support to cpy and enable

b10360

Model ReleasesDGX agent

common/peg : suppress incomplete escape sequences (#26780) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iO

BAG: Budget-Aware Gating for Diffusion Caching

Model ReleasesDGX agent

arXiv:2608.09231v1 Announce Type: new Abstract: Diffusion caching is a lightweight strategy that accelerates Diffusion Transformers (DiTs) by reusing intermediate features across denoising steps, but

Beyond Aggregate Calibration: Decomposing Income-Conditional Recall Disparities in Automated Credit Default Prediction

SafetyDGX agent

arXiv:2608.08202v1 Announce Type: new Abstract: Data-centric curation pipelines frequently rely on model confidence scores to flag and filter noisy or mislabeled training instances. Evaluating this fi

Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics

Model ReleasesDGX agent

arXiv:2512.05098v2 Announce Type: replace-cross Abstract: In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily targe

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training

SafetyDGX agent

arXiv:2608.09217v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform tas

Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction

Model ReleasesDGX agent

arXiv:2608.08459v1 Announce Type: cross Abstract: Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domain

Beyond the Capability Boundary: Zeroth-Order Optimization for Self-Evolving LLM Agents

Model ReleasesDGX agent

arXiv:2608.09292v1 Announce Type: cross Abstract: Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. H

Beyond the Node: Clade-level Selection for Efficient MCTS in Automatic Heuristic Design

TutorialsDGX agent

arXiv:2602.00549v2 Announce Type: replace Abstract: While Monte Carlo Tree Search (MCTS) shows promise in Large Language Model (LLM) based Automatic Heuristic Design (AHD), it suffers from a critical

Blumira launches Hearth, an AI command center that spans rival security tools

Model ReleasesDGX agent

Security operations platform startup Blumira Inc. today launched Hearth, a vendor-agnostic artificial intelligence command center that lets security teams investigate and act across their existing too

Bridging the Evaluation Gap: Standardized Benchmarks for Multi-Objective Search

Model ReleasesDGX agent

arXiv:2603.24084v2 Announce Type: replace Abstract: Empirical evaluation in multi-objective search (MOS) has historically suffered from fragmentation, relying on heterogeneous problem instances with i

Building Agent Skills and testing them is hard, but it doesn't have to be. Listen to Arjun Patel demo Cultivar, an open source tool develope…

Model ReleasesDGX agent

Building Agent Skills and testing them is hard, but it doesn't have to be. Listen to Arjun Patel demo Cultivar, an open source tool developed at Pinecone to help benchmark agent skills in sandboxes. T

ChatGPT and Gemini both just passed 1 billion users

Model ReleasesDGX agent

For the 14th time, a Google product has hit 1 billion users. Google CEO Sundar Pichai posted on X that a billion people are using Gemini every month, and that Gemini is Google's fastest-growing produc

ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons

Model ReleasesDGX agent

arXiv:2608.09282v1 Announce Type: new Abstract: Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise

Communication-efficient distributed hazard difference estimation for heterogeneous multi-site survival data

ResearchDGX agent

arXiv:2601.14609v2 Announce Type: replace-cross Abstract: Multi-site collaboration can power survival models that no single hospital could fit alone, but privacy rules and protected computing environm

Contextual Value Alignment via Multilayer Combinatorial Fusion

SafetyDGX agent

arXiv:2608.07642v1 Announce Type: new Abstract: Aligning large language models (LLMs) with human values remains a major challenge, especially for trustworthy AI. While existing approaches such as RLHF

Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation

ResearchDGX agent

arXiv:2608.09101v1 Announce Type: new Abstract: Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misali

Dancing Points: Synthesizing Ballroom Dancing with Three-Point Inputs

ResearchDGX agent

arXiv:2601.02096v2 Announce Type: replace-cross Abstract: Ballroom dancing is a structured yet expressive motion category. Its highly diverse movement and complex interactions between leader and follo

Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning

ResearchDGX agent

arXiv:2602.11149v2 Announce Type: replace Abstract: Supervised fine-tuning (SFT) on chain-of-thought data is an essential post-training step for reasoning language models. Standard machine learning in

Deep probabilistic logic programming for diagnostic reasoning from incomplete information: A case study in stroke detection

ApplicationsDGX agent

arXiv:2608.08561v1 Announce Type: new Abstract: In medical applications, raw data is frequently associated with significant privacy concerns, lending particular importance to the encoding of summary s

Deploying Anthropic Claude apps gateway for AWS for enterprise workloads

Model ReleasesDGX agent

Claude apps gateway is a self-hosted governance layer between Claude Code and Claude Desktop and Amazon Bedrock or Claude Platform on AWS. This post presents a production reference deployment covering

DH-VLM: Dual-Horizon Cooperative Latent Reasoning for Autonomous Driving

SafetyDGX agent

arXiv:2608.09333v1 Announce Type: new Abstract: Large-scale language models for autonomous driving enable enhanced global understanding and long-horizon planning. However, when deployed in isolated ve

Directional-Clamp PPO

Model ReleasesDGX agent

arXiv:2511.02577v2 Announce Type: replace Abstract: Proximal Policy Optimization (PPO) is widely regarded as one of the most successful deep reinforcement learning algorithms, known for its robustness

Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization

Model ReleasesDGX agent

arXiv:2608.09043v1 Announce Type: cross Abstract: Users of modern platforms repeatedly need summaries of recent dialogue, but the window rarely contains enough context to be interpreted on its own. We

DualCert: A Solver for the Traveling Salesman Problem with Constraint-Coupled Learning

Model ReleasesDGX agent

arXiv:2608.09042v1 Announce Type: new Abstract: Large traveling salesman problem (TSP) instances require a solver to allocate limited computation while preserving the validity of its outputs. Existing

EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference

HardwareDGX agent

arXiv:2608.07964v1 Announce Type: cross Abstract: Load Balancing has emerged as a critical problem in expert-parallel distributed inference of Mixture-of-Experts (MoE) models. As routing distributions

Efficient Fine-Tuning of DINOv3 Pretrained on Natural Images for Atypical Mitotic Figure Classification

Model ReleasesDGX agent

arXiv:2508.21041v4 Announce Type: replace-cross Abstract: Atypical mitotic figures (AMFs) indicate abnormal cell division associated with poor prognosis. Their detection remains difficult due to low p

Eikonal Regularisation in Physics-Informed Neural Networks for Three-Dimensional Level-Set Advection: Transferability of Two-Dimensional Design Principles

Model ReleasesDGX agent

arXiv:2608.08322v1 Announce Type: cross Abstract: Physics-informed neural networks applied to the level-set formulation of interface advection commonly augment the residual and initial-condition losse

Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation

ApplicationsDGX agent

arXiv:2510.21891v2 Announce Type: replace-cross Abstract: To deploy large language models (LLMs) in high-stakes application domains that require substantively accurate responses to open-ended prompts,

EndoMD-SLAM: Endoscopic Gaussian Splatting SLAM under Optical Degradation with Memory and Static-Transient Decomposition

Model ReleasesDGX agent

arXiv:2608.08949v1 Announce Type: new Abstract: Dense 3D reconstruction is critical for clinical endoscopic navigation and documentation. While Gaussian Splatting SLAM systems show promise in this dom

EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility

Model ReleasesDGX agent

arXiv:2608.08691v1 Announce Type: new Abstract: Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but physical flexibility becomes dependable capacity

Entropy-based Code Adversarial Translation for Real-world Repository Migration

Model ReleasesDGX agent

arXiv:2608.09273v1 Announce Type: new Abstract: LLMs have demonstrated strong capabilities in code generation and automated program repair, but migrating an entire repository rarely produces a runnabl

Exact Contraction Rates via the Berkson--Porta Representation: A Sharp Threshold and Its Herglotz-Kernel Obstruction

Model ReleasesDGX agent

arXiv:2608.07552v1 Announce Type: cross Abstract: Semigroups of holomorphic self-maps of the unit disc with an interior fixed point are, by the classical Berkson--Porta representation, entirely determ

Exploring LLM Capabilities for Situational Understanding and COLREG compliance on real-world maritime navigation scenarios

ApplicationsDGX agent

arXiv:2608.08281v1 Announce Type: new Abstract: Recently, Large Language Models (LLMs) have shown considerable capability for situational understanding, reasoning, and decision making in different dom

Failure-Mechanism Transferability of Cumulative-Damage Features for Health State Estimation of SiC Power Modules

Model ReleasesDGX agent

arXiv:2608.08365v1 Announce Type: cross Abstract: Data-driven health-state estimators for SiC (Silica-Carbide) power modules typically report their performance on a single accelerated-aging campaign,

Fair on the Surface? Benchmarking Hidden-Output Fairness Gaps in LLM Recommenders

Model ReleasesDGX agent

arXiv:2608.08284v1 Announce Type: new Abstract: Fairness audits for LLM-based recommenders have largely focused on observable outputs, implicitly assuming that stable recommendations reflect stable in

Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents

Model ReleasesDGX agent

arXiv:2608.08852v1 Announce Type: new Abstract: AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content. However, whether they can adapt a lesson to fi

Finite Constant Frontiers and Auditable Regret Certificates for Average-Reward Reinforcement Learning

Model ReleasesDGX agent

arXiv:2608.07725v1 Announce Type: new Abstract: Average-reward reinforcement-learning regret is known up to logarithmic factors, but the numerical content of published guarantees is difficult to compa

Floating-Point Neural Network Verification at the Software Level

Model ReleasesDGX agent

arXiv:2510.23389v2 Announce Type: replace-cross Abstract: The behaviour of neural network components must be proven correct before deployment in safety-critical systems. Unfortunately, existing neural

Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective

SafetyDGX agent

arXiv:2608.08445v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is widely regarded as a novel paradigm born from the limitations of large language models (LLMs)--a mechanism to gr

FriskAI launches with $3.6M to show enterprises what their AI agents are doing

Model ReleasesDGX agent

Runtime intelligence startup FriskAI Inc. launched today with 3.6 million in pre-seed funding to give enterprises a record of what their artificial intelligence agents actually do once they go into pr

From Product Search to Preference Articulation: The Economics of Agentic Commerce

Model ReleasesDGX agent

arXiv:2608.08395v1 Announce Type: cross Abstract: Generative AI is shifting digital commerce from browsing toward agentic search, in which consumers delegate product discovery to AI agents. We compare

From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability

ResearchDGX agent

arXiv:2608.08904v1 Announce Type: cross Abstract: How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action

From token probabilities to calibrated confidence: An empirical study of mathematical question answering

ResearchDGX agent

arXiv:2608.07827v1 Announce Type: cross Abstract: Confidence estimation for large language models (LLMs) aims to estimate the probability that a generated answer is correct, while calibration aligns t

FSTC-Encoder: Feature--Spatial--Temporal Correlation Learning for Generalizable RF Sensing

ResearchDGX agent

arXiv:2608.08439v1 Announce Type: cross Abstract: Heterogeneous RF sensing differs substantially in feature structure, spatial layout, and temporal scale, making existing models difficult to reuse acr

Google’s Gemini AI app passes 1 billion monthly active users

Model ReleasesDGX agent

Google LLC’s Gemini artificial intelligence app has passed 1 billion monthly active users, making it the 14th product in the company’s history to reach that mark. The company announced the milestone t

Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding

ResearchDGX agent

arXiv:2604.02047v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates large language model inference by drafting multiple candidate tokens and verifying them in a single forward p

GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning

Model ReleasesDGX agent

arXiv:2608.07905v1 Announce Type: new Abstract: Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to long-horizon tasks, and lack of environment

HeatCast: A Benchmark for Neighborhood-Scale LST Forecasting across 124 U.S. Cities

Model ReleasesDGX agent

arXiv:2608.07640v1 Announce Type: new Abstract: Land Surface Temperature (LST) is a widely used satellite-derived measure of urban surface heat, but there is no shared benchmark for forecasting it at

High Fidelity Capture, Reconstruction, and Transfer of Human Demonstrations for Robot-Assisted Bathing

Model ReleasesDGX agent

arXiv:2608.09127v1 Announce Type: new Abstract: Despite the demand for robots in high-value clinical tasks like bathing, contemporary systems still lack the safety and reliability required for complex

How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning

ResearchDGX agent

arXiv:2505.24273v2 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) suggest that reinforcement learning (RL) effectively internalizes search strategies, yielding si

Imaginative Generative AI: Crossing the Entropy Wall into Worlds Beyond Imitation

ResearchDGX agent

arXiv:2608.09385v1 Announce Type: cross Abstract: Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator

Implicit Virtual Leader: Decentralized Vision-Only Relative Pose Estimation for Multi-Robot Formations

Model ReleasesDGX agent

arXiv:2607.15708v2 Announce Type: replace Abstract: Classical leader-follower formation control suffers from single points of failure and error propagation, and relies on absolute localization sensors

← Previous
1…601602603604605…1060
Next →