AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “apple-ml-research”

GridTimelineEvolution
61+ results
7 Aug 2026

Arbitrage: Efficient Reasoning via Advantage-Aware Speculation

ResearchDGX agent

Modern Large Language Models achieve impressive reasoning capabilities with long Chain of Thoughts, but they incur substantial computational cost during inference, and this motivates techniques to imp

Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models

ResearchDGX agent

Large Language Models (LLMs) have achieved state-of-the-art performance on a broad range of Natural Language Processing (NLP) tasks, including document processing and code generation. Autoregressive L

Scaling Categorical Flow Maps

ResearchDGX agent

Continuous diffusion and flow matching models could represent a powerful alternative to autoregressive approaches for language modelling (LM), as they unlock a host of advantages currently reserved fo


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
6 Aug 2026

DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness

ResearchDGX agent

Large language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often struggle to produce complete answer set to complex questions such a

Locking Pretrained Weights via Deep Low-Rank Residual Distillation

ResearchDGX agent

The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling their use across diverse hardware and software plat

3 Aug 2026

Understanding Alignment in Multimodal LLMs: A Comprehensive Study

SafetyDGX agent

Preference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains comparatively under

28 Jul 2026

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

Local AiDGX agent

Siri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple’s most powerful on-device foundation model. This work presents th

27 Jul 2026

GH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision Tasks

ResearchDGX agent

Systematic failures of vision models on semantically coherent subsets, known as error slices, reveal limitations in robustness and evaluation. Existing slice discovery approaches largely model slices

24 Jul 2026

LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning

ResearchDGX agent

Long-horizon execution in Large Language Models (LLMs) remains unstable even when high-level strategies are provided. Evaluating on controlled algorithmic puzzles, we demonstrate that while decomposit

21 Jul 2026

Accelerating Text-to-Video Generation with Calibrated Sparse Attention

ResearchDGX agent

Recent diffusion models enable high-quality video generation, but suffer from slow runtimes. The large transformer-based backbones used in these models are bottlenecked by spatiotemporal attention. In

15 Jul 2026

CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning

ResearchDGX agent

Retrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge but still suffers from long contexts and disjoint retrieval–generation optimization. In this work, we

One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation

ResearchDGX agent

Visual generative models (e.g., diffusion models) typically operate in compressed latent spaces to balance training efficiency and sample quality. In parallel, there has been growing interest in lever

14 Jul 2026

Multilingual Semantic Retrieval for Apple Music Search

Model ReleasesDGX agent

Apple Music serves listeners across 150+ storefronts in dozens of languages, with a catalog that grows by hundreds of thousands of new tracks daily. At this scale, search recall on misspelled, transli

Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants

AgentsDGX agent

Proactive agents that anticipate user needs and autonomously execute tasks hold great promise as digital assistants, yet the lack of realistic user simulation frameworks hinders their development. Exi

10 Jul 2026

Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies

AgentsDGX agent

This paper was accepted at the AI4TCI (Workshop on AI for Secure and Trustworthy Critical Infrastructure Systems) Workshop at the International Conference on Availability, Reliability and Security (AR

9 Jul 2026

Incentivizing Temporal-Awareness in Egocentric Video Understanding Models

SafetyDGX agent

Multimodal large language models (MLLMs) have recently shown strong performance in visual understanding, yet they often lack temporal awareness, particularly in egocentric settings where reasoning dep

Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context

AgentsDGX agent

Long-context handling remains a core challenge for language models: even with extended context windows, models often fail to reliably extract, reason over, and use the information across long contexts

Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why

SafetyDGX agent

On-policy distillation offers dense, per-token supervision for training reasoning models; however, it remains unclear under which conditions this signal is beneficial and under which it is detrimental

7 Jul 2026

FlowEval: Reference-Based Evaluation of Generated User Interfaces

ResearchDGX agent

While large language models (LLMs) and coding agents are often applied to user interface (UI) development, developers find it difficult to reliably assess their proficiency in visual and interaction d

MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching

SafetyDGX agent

Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality required by everyday

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

ResearchDGX agent

This study focuses on Text-to-Sounding-Video (T2SV) generation, which aims to generate a video with synchronized audio from text, with both modalities aligned to the text conditions. Despite progress

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents

ResearchDGX agent

The web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to offline trajectories fo

6 Jul 2026

Path-Constrained Mixture-of-Experts

ResearchDGX agent

Sparse Mixture-of-Experts (MoE) architectures route each token through a subset of experts at each layer independently. We propose viewing MoE computation through the lens of expert paths—the sequence

Revisiting ASR Error Correction with Specialized Models

ResearchDGX agent

Language models play a central role in automatic speech recognition (ASR), yet most methods rely on text-only models unaware of ASR error patterns. Recently, large language models (LLMs) have been app

Segmental Attention Decoding with Long Form Acoustic Encodings

TutorialsDGX agent

We address the fundamental incompatibility of attention-based encoder-decoder (AED) models with long-form acoustic encodings. AED models trained on segmented utterances learn to encode absolute frame

Understanding Annotator Safety Policy with Interpretability

SafetyDGX agent

Safety policies define what constitutes safe and unsafe AI outputs, guiding data annotation and model development. However, annotation disagreement is pervasive and can stem from multiple sources such

2 Jul 2026

Amortizing Maximum Inner Product Search with Learned Support Functions

ResearchDGX agent

Maximum inner product search (MIPS) is a crucial subroutine in machine learning, requiring the identification of a vector taken within a database (the keys) that best aligns with a given query. We pro

Anti-Causal Domain Generalization: Leveraging Unlabeled Data

ResearchDGX agent

The problem of domain generalization concerns learning predictive models that are robust to distribution shifts when deployed in new, previously unseen environments. Existing methods typically require

Learning Structured Reasoning via Tractable Trajectory Control

ResearchDGX agent

Large language models can exhibit emergent reasoning behaviors, often manifested as recurring lexical patterns (e.g., “wait,” indicating verification). However, complex reasoning trajectories remain s

Learning Unmasking Policies for Diffusion Language Models

ResearchDGX agent

Diffusion (Large) Language Models (dLLMs) now match the downstream performance of their autoregressive counterparts on many tasks, while holding the promise of being more efficient during inference. O

MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers

ResearchDGX agent

Understanding how transformer components operate in LLMs is important, as it is at the core of recent technological advances in artificial intelligence. In this work, we revisit the challenges associa

Multi-Agent Teams Hold Experts Back

AgentsDGX agent

Multi-agent LLM systems are increasingly deployed as autonomous collaborators, where agents interact freely rather than execute fixed, pre-specified workflows. In such settings, effective coordination

On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs

ResearchDGX agent

Reinforcement learning (RL) finetuning has become a key technique for enhancing large language models (LLMs) on reasoning-intensive tasks, motivating its extension to vision language models (VLMs). Wh

Residual Context Diffusion Language Models

ResearchDGX agent

Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to purely autoregressive language models because they can decode multiple tokens in parallel. However, state-of-the-art

23 Jun 2026

Metric-Dependent Annotation Saturation for Learning from Label Distributions

ResearchDGX agent

When annotators disagree on a label, the disagreement itself carries signal—and the number of annotators needed to capture it depends on the evaluation metric. We fine-tune NLI models on label distrib

8 Jun 2026

Introducing the Third Generation of Apple’s Foundation Models

Local AiDGX agent

Our next generation of Apple Intelligence is centered around our users, integrated deeply into our operating systems, and powered by a bold new architecture with privacy at its core. At the heart of t

28 May 2026

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

ResearchDGX agent

CVPR 2026 is a major international conference on computer vision and pattern recognition organized by IEEE and the Computer Vision Foundation, where Apple presents its latest machine learning research

22 May 2026

VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models

ResearchDGX agent

Streaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core mechanism for real-time visual assistants. Exis

8 May 2026

Apple Workshop on Privacy-Preserving Machine Learning & AI 2026

ResearchDGX agent

At Apple, we believe privacy is a fundamental human right. As AI capabilities increase and become more integrated into people’s daily lives, advancing research in privacy-preserving techniques is incr

RVPO: Risk-Sensitive Alignment via Variance Regularization

SafetyDGX agent

Current critic-less RLHF methods aggregate multi-objective rewards via an arithmetic mean, leaving them vulnerable to constraint neglect: high-magnitude success in one objective can numerically offset

Velox: Learning Representations of 4D Geometry and Appearance

ResearchDGX agent

We introduce a framework for learning latent representations of 4D objects which are descriptive, faithfully capturing object geometry and appearance; compressive, aiding in downstream efficiency; and

7 May 2026

What Matters in Practical Learned Image Compression

ResearchDGX agent

One of the major differentiators unlocked by learned codecs relative to their hard-coded traditional counterparts is their ability to be optimized directly to appeal to the human visual system. Despit

6 May 2026

From Where Things Are to What They’re For: Benchmarking Spatial–Functional Intelligence for Multimodal LLMs

Model ReleasesDGX agent

True spatial intelligence for multimodal agents transcends low-level geometric perception, evolving from knowing where things are to understanding what they are for. While existing benchmarks, such as

SpecMD: A Comprehensive Study on Speculative Expert Prefetching

ResearchDGX agent

Mixture-of-Experts (MoE) models enable sparse expert activation, meaning that only a subset of the model’s parameters is used during each inference. However, to translate this sparsity into practical

30 Apr 2026

Bootstrapping Sign Language Annotations with Sign Language Models

ResearchDGX agent

AI-driven sign language interpretation is limited by a lack of high-quality annotated data. New datasets including ASL STEM Wiki and FLEURS-ASL contain professional interpreters and 100s of hours of d

International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2026

ResearchDGX agent

ICASSP 2026 is a major international conference focused on acoustics, speech, and signal processing research and applications. Apple's machine learning research team is participating in the conference

STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows

ResearchDGX agent

Normalizing flows (NFs) are end-to-end likelihood-based generative models for continuous data, and have recently regained attention with encouraging progress on image generation. Yet in the video gene

29 Apr 2026

Adaptive Thinking: Large Language Models Know When to Think in Latent Space

ResearchDGX agent

Recent advances in large language models (LLMs) test-time computing have introduced the capability to perform intermediate chain-of-thought (CoT) reasoning (thinking) before generating answers. While

DSO: Direct Steering Optimization for Bias Mitigation

SafetyDGX agent

Generative models are often deployed to make decisions on behalf of users, such as vision-language models (VLMs) identifying which person in a room is a doctor to help visually impaired individuals. Y

28 Apr 2026

LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning

ResearchDGX agent

Large Language Models (LLMs) demonstrate their reasoning ability through chain-of-thought (CoT) generation. However, LLM’s autoregressive decoding may limit the ability to revisit and refine earlier t

Local Mechanisms of Compositional Generalization in Conditional Diffusion

TutorialsDGX agent

Conditional diffusion models appear capable of compositional generalization, i.e., generating convincing samples for out-of-distribution combinations of conditioners, but the mechanisms underlying thi

StereoFoley: Object-Aware Stereo Audio Generation from Video

ResearchDGX agent

We present StereoFoley, a video-to-audio generation framework that produces semantically aligned, temporally synchronized, and spatially accurate stereo sound at 48 kHz. While recent generative video-

24 Apr 2026

Learning Long-Term Motion Embeddings for Efficient Kinematics Generation

ResearchDGX agent

Understanding and predicting motion is a fundamental component of visual intelligence. Although modern video models exhibit strong comprehension of scene dynamics, exploring multiple possible futures

23 Apr 2026

ParaRNN: Large-Scale Nonlinear RNNs, Trainable in Parallel

ResearchDGX agent

Recurrent Neural Networks (RNNs) are naturally suited to efficient inference, requiring far less memory and compute than attention-based architectures, but the sequential nature of their computation h

22 Apr 2026

Apple Machine Learning Research at ICLR 2026

ResearchDGX agent

Apple is advancing AI and ML with fundamental research, much of which is shared through publications and engagement at conferences in order to accelerate progress in this important field and support t

21 Apr 2026

Can Large Language Models Understand Context?

Model ReleasesDGX agent

Understanding context is key to understanding human language, an ability which Large Language Models (LLMs) have been increasingly seen to demonstrate to an impressive extent. However, though the eval

17 Apr 2026

International Conference on Learning Representations (ICLR) 2026

ResearchDGX agent

ICLR 2026 is a major international conference on machine learning and deep learning where Apple's research team will present their latest work and innovations in representation learning. The conferenc

16 Apr 2026

MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining

ResearchDGX agent

This paper was accepted at the Workshop on Navigating and Addressing Data Problems for Foundation Models (NADPFM) at ICLR 2026. Principled domain reweighting can substantially improve sample efficienc

10 Apr 2026

ACM Human-Computer Interaction Conference (CHI) 2026

ResearchDGX agent

Apple is presenting new research at the ACM CHI Conference on Human Factors in Computing Systems, held in person in Barcelona, Spain, from April 13–17, 2026, and is a sponsor of the event. Research...

9 Apr 2026

A Theoretical Framework for Acoustic Neighbor Embeddings

ResearchDGX agent

This paper provides a theoretical framework for interpreting acoustic neighbor embeddings, which are representations of the phonetic content of variable-width audio or text in a fixed-dimensional embe

← Previous
1
Next →
62 results
← Previous
12
Next →