AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,522 results
7 Aug 2026

Toward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based Model

Model ReleasesDGX agent

arXiv:2608.06252v1 Announce Type: cross Abstract: Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL). Automatic BdSL recognition on personal devices c

Uncertainty-Aware World Model for Aerial Image-Goal Navigation

ApplicationsDGX agent

arXiv:2608.05597v1 Announce Type: new Abstract: Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

SafetyDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.05799v1 Announce Type: cross Abstract: Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to

6 Aug 2026

AMA: MiniMax H3 Team — Ask us anything about our open video generation model, training, and future plans

Local AiDGX agent

https://preview.redd.it/kihat320ashh1.png?width=1672&format=png&auto=webp&s=a7ccc40ba3fb229ac7ebf57e8e6a314e0ee45646 Hi r/StableDiffusion! u/New-Requirement1419 -> dacongya (Head of H3 Researcher) u/A

Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation

ResearchDGX agent

arXiv:2511.09002v3 Announce Type: replace-cross Abstract: Self-consuming generative models have received significant attention over the last few years. In this paper, we study a self-consuming generat

Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations

ResearchDGX agent

arXiv:2604.00209v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in high-stakes settings, yet they frequently violate contextual privacy by disclosing private

EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks

Model ReleasesDGX agent

arXiv:2608.04549v1 Announce Type: cross Abstract: Frontier LLMs are increasingly put to use on open-ended complex questions, different in nature from the ones they are typically evaluated on. We dedic

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning

ResearchDGX agent

arXiv:2608.04646v1 Announce Type: new Abstract: Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and validity of the

Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models

ResearchDGX agent

arXiv:2608.04765v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic control. However,

FLUX 3 is now live on Together AI. @bfl_ai’s new multimodal model generates video and synchronized audio together, with up to 20-second clip…

ToolsDGX agent

FLUX 3 is now live on Together AI. @bfl_ai’s new multimodal model generates video and synchronized audio together, with up to 20-second clips, multiple shots, and control from text, images, or keyfram

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation

Model ReleasesDGX agent

arXiv:2608.04378v1 Announce Type: cross Abstract: Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow

Koopman-Based Nonlinear Identification and Model Predictive Control of a Turbofan Engine

ResearchDGX agent

arXiv:2604.01730v2 Announce Type: replace Abstract: This paper investigates Koopman operator-based approaches for multivariable control of a two-spool turbofan engine. A physics-based component-level

LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content

Model ReleasesDGX agent

arXiv:2410.10783v4 Announce Type: replace Abstract: The large-scale training of multi-modal models on data scraped from the web has shown outstanding utility in infusing these models with the required

MarsCast: Transfer Learning of AI Weather Foundation Models to Planetary Atmospheres

ResearchDGX agent

arXiv:2608.05054v1 Announce Type: cross Abstract: We investigate the transferability of Earth weather foundation models to planetary atmospheres by adapting the GraphCast graph neural weather forecast

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2608.04633v1 Announce Type: new Abstract: Recent Vision-Language-Action (VLA) methods improve generalization by aligning their representations with 3D scene geometry. However, these methods are

The death of SLMs?

Model ReleasesDGX agent

I love to see these impressive models coming out that compete with the giants from companies like Z.ai, Moonshot, Alibaba, etc. A win for the open source/weight community is always welcome. While I am

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

SafetyDGX agent

arXiv:2608.04436v1 Announce Type: new Abstract: Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understandi

When Absence Is Evidence: Evaluating Completeness-Sensitive Negative Reasoning in Large Language Models

ResearchDGX agent

arXiv:2608.04591v1 Announce Type: cross Abstract: Large language models (LLMs) are often asked whether something is absent from a record, list, or retrieved context. Yet non-observation licenses a neg

5 Aug 2026

Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging

Local AiDGX agent

arXiv:2608.03316v1 Announce Type: cross Abstract: On-policy distillation, in which a teacher corrects samples that the student itself generates, presupposes that the two models speak the same language

ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2608.03358v1 Announce Type: new Abstract: Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emot

BOW: Training Language Models to Reason Over Plausible Next Words

SafetyDGX agent

arXiv:2506.13502v3 Announce Type: replace Abstract: Next-word prediction (NWP) trains language models against a single observed continuation, even though many contexts admit multiple plausible next wo

Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss?

Model ReleasesDGX agent

arXiv:2608.03983v1 Announce Type: cross Abstract: Optimizing compilers miss profitable transformations when their enabling semantics are absent from the analyzed program representation. We ask whether

Can Training Logs Make Model Comparisons More Precise?

ResearchDGX agent

arXiv:2608.02705v1 Announce Type: cross Abstract: Comparing stochastically trained models requires estimating both a performance difference and its uncertainty from repeated runs. We study whether tra

Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse

ApplicationsDGX agent

arXiv:2608.03893v1 Announce Type: new Abstract: Production deployments often swap between different-sized models in a family for cost-quality cascading, mid-conversation switching, and routing, and ea

Cura 1T: Specialized Model for Agentic Healthcare

Model ReleasesDGX agent

arXiv:2607.15314v2 Announce Type: replace Abstract: Healthcare AI agents handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR)

DUD: Decoupled Update Dynamics for Reliable Uncertainty Quantification in Large Language Models

ResearchDGX agent

arXiv:2608.03411v1 Announce Type: new Abstract: Accurate Uncertainty Quantification (UQ) is critical for reliable deployment of Large Language Models (LLMs), yet traditional probability-based metrics

Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement

ResearchDGX agent

arXiv:2608.03733v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-

Mapping the City Through the Lens of Language Models

ResearchDGX agent

arXiv:2608.02971v1 Announce Type: new Abstract: Language models often complete an underspecified reference to a city with unstated assumptions about urban size, form, infrastructure, environment, and

Noise-Aware Shrinkage for Differentially Private Zeroth-Order Fine-Tuning of Large Language Models

ResearchDGX agent

arXiv:2608.03277v1 Announce Type: new Abstract: Differentially private zeroth-order optimization (DP-ZO) enables memory-efficient private fine-tuning of large language models using only forward evalua

'OCR is just a feature now. Frontier models will eat it.' We hear this constantly. The data says otherwise. Across three GPT generations, pa…

AgentsDGX agent

'OCR is just a feature now. Frontier models will eat it.' We hear this constantly. The data says otherwise. Across three GPT generations, parsing accuracy gained ~24 points, while cost per page 4x'd.

Scaling agentic AI: How UiPath built its high-performance GPU platform on AI Hypercomputer

Model ReleasesDGX agent

As a market leader in enterprise agentic automation and business orchestration, UiPath is helping to pioneer an industry shift toward agentic AI. With it, the company is deploying autonomous agents to

Speculative Correction: Draft-then-Refine Decoding for Diffusion Language Models

Local AiDGX agent

arXiv:2608.02625v1 Announce Type: cross Abstract: Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right generation by p

T2VAttack: Adversarial Attack on Text-to-Video Diffusion Models

SafetyDGX agent

arXiv:2512.23953v2 Announce Type: replace Abstract: The rapid evolution of Text-to-Video (T2V) diffusion models has driven remarkable advancements in generating high-quality, temporally coherent video

Where Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models

ApplicationsDGX agent

arXiv:2601.09445v2 Announce Type: replace-cross Abstract: In language models (LMs), intra-memory knowledge conflict arises when inconsistent information about the same subject is encoded within the mo

4 Aug 2026

Development and Validation of a Dynamic Kidney Failure Prediction Model based on Deep Learning: A Real-World Study with External Validation

ApplicationsDGX agent

arXiv:2501.16388v3 Announce Type: replace Abstract: Background: Chronic kidney disease (CKD), a progressive disease with high morbidity and mortality, has become a significant global public health pro

Douyin Multimodal Embedding Model Technical Report

ApplicationsDGX agent

arXiv:2608.02148v1 Announce Type: cross Abstract: Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search

Faster-WAM: Do World Action Models Need Deep Action Modules?

ResearchDGX agent

arXiv:2608.02365v1 Announce Type: cross Abstract: World Action Models (WAMs) couple robot action prediction with video world models. Existing WAMs with shared-backbone and Mixture-of-Transformers desi

Few-Shot Biomedical Relation Extraction with Large Language Models: A Viable Alternative to Supervised Learning?

ResearchDGX agent

arXiv:2606.15412v2 Announce Type: replace Abstract: Biomedical relation extraction (BioRE) is a key step in transforming biomedical literature into structured knowledge. Most existing approaches rely

Isotonic Bradley-Terry Model for Paired Comparison Data

TutorialsDGX agent

arXiv:2608.02081v1 Announce Type: new Abstract: In this paper, we study prediction problems for paired comparison data, for example, predicting the win probability between two unmatched players and ra

OpenAI says one of its models exploited a website after third-party AI security lab Irregular mistakenly gave it access to the internet during evaluations (Wired)

AgentsDGX agent

Wired: OpenAI says one of its models exploited a website after third-party AI security lab Irregular mistakenly gave it access to the internet during evaluations — Rogue AI agents from OpenAI and Anth

Orchestrating Dual-Boundaries: An Arithmetic Intensity Inspired Acceleration Framework for Diffusion Language Models

ResearchDGX agent

arXiv:2511.21759v2 Announce Type: replace Abstract: Diffusion-based large language models (dLLMs) have recently gained significant attention for their exceptional performance and inherent potential fo

PlantRig - From Bones to Branches: Adaptation of Autoregressive Rigging Models for Plant Skeletal Reconstruction

SafetyDGX agent

arXiv:2608.01072v1 Announce Type: new Abstract: Autoregressive rigging models such as UniRig and SkinTokens perform well on articulated characters, but their ability to generalize to plant structures

Representation Transfer of Foundation Models for Ultra-Widefield Retinal Imaging

ResearchDGX agent

arXiv:2608.00586v1 Announce Type: new Abstract: Despite the widespread adoption of foundation models as feature extractors for medical imaging, relatively little is understood about how different pret

Retrieval-Augmented Interpretable Learning: Towards Task-Specific Zero-Shot Models in Healthcare

ApplicationsDGX agent

arXiv:2607.17508v2 Announce Type: replace Abstract: We introduce Retrieval-Augmented Interpretable Learning (RAIL), a probabilistic meta-learning framework for zero-shot generation of task-specific in

Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments

SafetyDGX agent

arXiv:2608.00419v1 Announce Type: cross Abstract: Large language models deployed in real-time, regulated settings face knowledge staleness, catastrophic forgetting, hallucination, and weak feedback lo

VARPose: Flexible 2D Pose Densification via Visual Autoregressive Modeling for Enhanced 3D Lifting

ResearchDGX agent

arXiv:2608.02214v1 Announce Type: new Abstract: Visual AutoRegressive Modeling (VAR) has excelled in natural image generation via next-scale prediction, but its use on topology-structured data like hu

Why are Chinese models better* at Frontend than the western top labs?

Model ReleasesDGX agent

I use A LOT both openAI and Anthropic products. When I need some frontend work (pure web dev) (or answer that feel less verbose and more to the point) I use Anthropic. For multimodality openAI feels b

3 Aug 2026

ActionParty: Multi-Subject Action Binding in Generative Video Games

Model ReleasesDGX agent

arXiv:2604.02330v2 Announce Type: replace-cross Abstract: Recent advances in video diffusion have enabled the development of 'world models' capable of simulating interactive environments. However, the

An analysis of machine learning approaches for enhancing decision-making in complex discrete choice tasks

SafetyDGX agent

arXiv:2607.28854v1 Announce Type: new Abstract: Discrete choice modeling is a common tool used for preference elicitation during policy-making, but this is typically done through parametric models. Ma

AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models

ResearchDGX agent

arXiv:2505.20255v3 Announce Type: replace Abstract: Recent advances in video diffusion models have substantially enhanced character animation techniques. However, existing methods primarily depend on

Beyond Feature and Structure Alignment: Learning Transferable Propagation Knowledge for Graph Foundation Models

SafetyDGX agent

arXiv:2607.28980v1 Announce Type: new Abstract: Graph Foundation Models (GFMs) have recently emerged as a promising paradigm for enabling knowledge transfer across diverse domains. Unlike traditional

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

Model ReleasesDGX agent

arXiv:2607.28631v1 Announce Type: new Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and compari

China’s Alibaba takes another swipe at America’s AI supremacy

Model ReleasesDGX agent

Chinese tech giant Alibaba released what it says is its largest and 'most capable AI model to date,' claiming performance rivaling the best systems from US frontier labs Anthropic and OpenAI, as well

Hugging Face CEO Clément Delangue weighs in on the open-weight AI models debate and said it helped the company when it was hacked by an Open…

IndustryDGX agent

Clément Delangue, CEO of Hugging Face, said the company’s focus on open‑weight AI models helped it withstand a hacking incident involving an unreleased OpenAI model. He also noted that China is leadin

On the Efficacy of Self-Supervised Point Cloud Encoders for Efficient 3D Large Language Models

SafetyDGX agent

arXiv:2607.29136v1 Announce Type: new Abstract: 3D point cloud-language models (3D-LLMs) enable 3D understanding by pairing point cloud encoders with large language models, but existing methods rely o

QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models

SafetyDGX agent

arXiv:2607.29445v1 Announce Type: cross Abstract: Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual question answering

The Morphological Core of Dungan: A Two-Dialect Finite-State Model and a Multi-Genre Evaluation

Model ReleasesDGX agent

arXiv:2607.28766v1 Announce Type: new Abstract: Dungan, a Sinitic language of Central Asia written in a Cyrillic-based script, is described in detail in the grammatical literature, yet the quantitativ

2 Aug 2026

Encrypted Clouds?

Model ReleasesDGX agent

I love the progress happening on open models but I feel like it is kind of getting clear that hardware to run good sized models is completely unaffordable for me right now. I know that you all love Qw

31 Jul 2026

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

ResearchDGX agent

arXiv:2607.28617v1 Announce Type: cross Abstract: System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout com

AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans

TutorialsDGX agent

arXiv:2607.28487v1 Announce Type: new Abstract: Fine-grained segmentation of auricular structures in CT is challenging because the ear occupies a small image region, cartilage boundaries are highly ir

← Previous
1…124125126127128…1009
Next →