microsoft/VibeVoice
microsoft/VibeVoice VibeVoice is Microsoft's Whisper-style audio model for speech-to-text, MIT licensed and with speaker diarization built into the model. Microsoft released it on January 21st, 2026 b
Knowledge catalogue
microsoft/VibeVoice VibeVoice is Microsoft's Whisper-style audio model for speech-to-text, MIT licensed and with speaker diarization built into the model. Microsoft released it on January 21st, 2026 b
🗓️ @mondaydotcom is scheduled for Interrupt. Monday Sidekick is an AI assistant that can execute work, manage tasks, and conduct research autonomously. At Interrupt, the Agent Conference by LangChain,
arXiv:2410.21548v3 Announce Type: replace Abstract: Large language models have drastically changed the prospects of AI by introducing technologies for more complex natural language processing. However
Pay attention to this one, AI devs. If you're building multi-agent systems, you're probably wiring static org charts. New research argues they should look more like a labor market. The paper introduce
arXiv:2604.22157v1 Announce Type: cross Abstract: Existing research typically treats privacy policies as flat, uniform text, extracting information without regard for the document's logical hierarchy.
Artificial intelligence may be dominating boardroom agendas, but many enterprises are discovering that the biggest obstacle to meaningful adoption is the state of their data. While consumer-facing AI
arXiv:2604.22746v1 Announce Type: cross Abstract: ReLU neural networks trained as surrogate models can be embedded exactly in mixed-integer linear programs (MILPs), enabling global optimization over t
arXiv:2509.14127v2 Announce Type: replace Abstract: We consider the problem of delivering multiple packages from a single depot to distinct goal locations using a homogeneous fleet of robots with limi
arXiv:2601.14541v4 Announce Type: replace-cross Abstract: This report distills the discussions and recommendations from the NSF Workshop on AI for Electronic Design Automation (EDA), held on December
arXiv:2604.22156v1 Announce Type: cross Abstract: Purpose: Accurate assessment of the Critical View of Safety (CVS) during laparoscopic cholecystectomy is essential to prevent bile duct injury, a comp
Gary Marcus critiques the marketing of large language model autocomplete systems as 'artificial intelligence,' arguing this terminology misrepresents their actual capabilities. He extends this critici
arXiv:2604.22351v1 Announce Type: cross Abstract: Mid-infrared astronomy from the ground faces critical challenges in accurately detecting and quantifying sources due to the dominant spatially and tim
Paul Graham's essay explores parenting philosophy and child-rearing principles, likely drawing on his experiences as a parent and his broader perspectives on education and human development. Sam Altma
TuneForge is an MCP (Model Context Protocol) server that enables coding agents like Claude and Cursor to perform machine learning operations directly within chat interfaces, including dataset generati
AUTOMATIC1111 remains the most popular choice for most users , though ComfyUI offers node-based workflows with more control and faster performance, while AUTOMATIC1111 provides a traditional interface
extremely happy that we are in q2 2026, and engineers i look up to are plowing a path for the local model future. yes, we are still beholden to some lab publishing weights. but i take that over compan
AURA is a local-first image management application designed specifically for AI enthusiasts and users who collect large numbers of AI-generated images. It features integration with Civitai, automated
Sam Altman commented on a discussion about AI efficiency, noting that a particular capability compresses work that previously took approximately two weeks into a single day, indicating significant pro
Next up: Hermes Agents on Tuesday in Lisbon 🇵🇹 as we work to make our office the go-to space for AI builders in Portugal, then we have another major one this week. we will this time be joined by the t
This is true only if you start after 9/11, exclude all BLM rioting, and count all killings by white prison gangs - but not minority ones like the Black Guerilla Family and La Raza ('The Race') groups.
This Reddit post from r/StableDiffusion asks the community for recommendations on which prompt engineering techniques and AI image generation models (specifically comparing z-image and z-image turbo v
Jamie Tarabay / Bloomberg: A profile of Strider, an intelligence firm that leverages agentic AI and public records to help the US Air Force, NATO, and others identify foreign state actors — Trump's ec
Alpha Eval: Agents Making Evals as a Multi-Player Game This diagram and blurb is largely a research riff with Claude on building data generation systems to get closer to the holy grail of self-improvi
Enterprises are rapidly moving from an artificial intelligence that answers questions and generates content to one that performs tasks and takes actions. According to Google Cloud Chief Executive Thom
Great paper on improving proactive agents. (bookmark it) Proactive agents act before you do. But how do you evaluate something that's supposed to anticipate needs you haven't expressed? This work intr
Groups and movements that can build & get implemented clear policies will have an outsized impact on the chances that AI is used in the way that they want. This is especially true in the near term It
Elon Musk announced a new version of Grok's Imagine model with improved lip synchronization and audio capabilities for AI-generated video content. The announcement emphasizes that all content created
ComfyUI announced a $500 million valuation in a notable post on X (formerly Twitter). The post appears to reference a specific or humorous way of announcing this significant milestone, likely celebrat
This is a useful image for thinking about the curve we are on and what likely comes next in an intuitively understandable way. there will this brief era where we can watch our AIs bumble around on the
arXiv:2509.03294v3 Announce Type: replace-cross Abstract: The increasing availability of personal data has enabled significant advances in fields such as machine learning, healthcare, and cybersecurit
arXiv:2604.21030v1 Announce Type: cross Abstract: The integration of Model Predictive Control (MPC) and Reinforcement Learning (RL) has emerged as a promising paradigm for constrained decision-making
arXiv:2604.21527v1 Announce Type: new Abstract: Low-cost air quality sensors (LCS) provide a practical alternative to expensive regulatory-grade instruments, making dense urban monitoring networks pos
arXiv:2604.21725v1 Announce Type: cross Abstract: LLM agents increasingly operate in open-ended environments spanning hundreds of sequential episodes, yet they remain largely stateless: each task is s
arXiv:2604.21744v1 Announce Type: cross Abstract: The capabilities of AI-assisted coding are progressing at breakneck speed. Chat-based vibe coding has evolved into fully fledged AI-assisted, agentic
arXiv:2511.23159v2 Announce Type: replace-cross Abstract: Vibe coding, the much-touted use of AI techniques for programming, faces two overwhelming obstacles: the difficulty of specifying goals ('prom
B8924 is a release build number from the llama.cpp project, a C/C++ framework for efficient large language model inference on consumer hardware. The llama.cpp project publishes multiple releases in a
arXiv:2604.21508v1 Announce Type: new Abstract: Protein-ligand bioactivity data published in the literature are essential for drug discovery, yet manual curation struggles to keep pace with rapidly gr
arXiv:2604.21854v1 Announce Type: new Abstract: Artificial intelligence now decides who receives a loan, who is flagged for criminal investigation, and whether an autonomous vehicle brakes in time. Go
Reuters: Brazil's Finance Minister Dario Durigan says the country has blocked prediction market platforms and tightened derivatives rules to curb “bet-like” products — Brazil has blocked prediction ma
In this post, we show how connecting the Visier Workforce AI platform with Amazon Quick through Model Context Protocol (MCP) gives every knowledge worker a unified agentic workspace to ask questions i
OpenRouter released a guide focusing on optimizing the use of Hermes Agent through the OpenRouter platform. The guide likely covers best practices, configuration options, and techniques for maximizing
arXiv:2603.18677v2 Announce Type: replace-cross Abstract: Artificial intelligence is increasingly embedded in human decision making. In some cases, it enhances human reasoning. In others, it fosters e
ComfyUI announced $17 million in funding from investors including Pace Capital, Chemistry, and Abstract Ventures . The funding will be used to scale Comfy Cloud, improve the local user experience, and
arXiv:2509.21361v2 Announce Type: replace-cross Abstract: Large language model (LLM) providers boast big numbers for maximum context window sizes. To test the real world use of context windows, we 1)
Deep Agents middleware is a great invention. You get a really strong default harness and a clean way to customize every part of it DeepAgents - it’s all you need - batteries included if you want, but
arXiv:2502.03484v2 Announce Type: replace-cross Abstract: Dementia encompasses a group of syndromes that impair cognitive functions such as memory, reasoning, and the ability to perform daily activiti
arXiv:2604.21187v1 Announce Type: cross Abstract: Ramsey-good graphs are graphs that contain neither a clique of size s nor an independent set of size t. We study doubly saturated Ramsey-good graphs,
arXiv:2506.19579v3 Announce Type: replace-cross Abstract: Robotic scene understanding increasingly relies on Vision-Language Models (VLMs) to generate natural language descriptions of the environment.
arXiv:2506.09998v2 Announce Type: replace-cross Abstract: Large language models (LLMs) can often accurately describe probability distributions using natural language, yet they still struggle to genera
Grok Voice is used by @Starlink Introducing Grok Voice Think Fast 1.0 A state-of-the-art voice model built for complex, multi-step workflows with snappy responses and high accuracy. It takes the top s
A user asks Harrison Chase about whether there are deepagents skills comparable to the claude-api skill that could help with building applications using deepagents framework. This appears to be a ques
Introducing Hermes Agent v0.11.0 Our largest update yet, with over 700 PRs across ~200 contributors. Thank you to everyone who's worked on Hermes Agent! This update features a beta TUI v2, unlimited r
I'd need to search for this specific Reddit discussion to provide an accurate summary of what was actually discussed. Let me retrieve that information. This Reddit post discusses using AI vision model
arXiv:2604.21870v1 Announce Type: cross Abstract: STEM education researchers are often interested in identifying moments of students' mechanistic reasoning for deeper analysis, but have limited capaci
arXiv:2604.21092v1 Announce Type: new Abstract: Integrating Large Language Models (LLMs) into complex software systems enables the generation of human-understandable explanations of opaque AI processe
arXiv:2601.20706v2 Announce Type: replace-cross Abstract: Diffusion-based LLMs (dLLMs) fundamentally depart from traditional autoregressive (AR) LLM inference: they leverage bidirectional attention, b
arXiv:2602.01493v2 Announce Type: replace-cross Abstract: Solving diverse partial differential equations (PDEs) is fundamental in science and engineering. Large language models (LLMs) have demonstrate
Our teams have been busyyy! Here are some key updates from the past week: — @GoogleCloud unveiled a suite of AI innovations at our Cloud Next event, including our eighth generation TPUs (TPUt for infe
arXiv:2604.21216v1 Announce Type: cross Abstract: The First Fundamental Theorem of Welfare Economics assumes that welfare-bearing agents are autonomous and implicitly relies on a binary distinction be
arXiv:2505.11702v2 Announce Type: replace Abstract: This work develops a framework for post-training augmentation invariance, in which our goal is to add invariance properties to a pretrained network