Research
Hosting Live session for sub 10ms retrieval by Moss (YC backed) [N]
Moss is a YC-backed high-performance runtime for real-time semantic search that delivers sub-10ms lookups, instant index updates, and zero infrastructure overhead, running where the agent lives — clou
Moss is a YC-backed high-performance runtime for real-time semantic search that delivers sub-10ms lookups, instant index updates, and zero infrastructure overhead, running where the agent lives — cloud, in-browser, or on-device. The company is working closely with Voice AI orchestration platforms like Pipecat and LiveKit, embedding Moss at the core of their real-time retrieval and context pipelines. The r/MachineLearning post promotes a live session focused on demonstrating and discussing how Moss achieves its sub-10ms semantic retrieval capabilities, likely targeting ML engineers and developers building conversational or voice AI applications.
Related
- What if your HNSW index stored 3-bit embeddings instead of float32? [R]
- Anyone have an S3-compatible store that actually saturates H100s without the AWS egress tax? [R]
- AI Systems Performance Engineering by Chris Fregly - is it worth it? [D]
- [[d-60-matmul-performance-bug-in-cublas-on-rtx-5090-d|[D] 60% MatMul Performance Bug in cuBLAS on RTX 5090 [D]]]
Source: r/MachineLearning | 2026-04-15