Hide to Guide: Learning via Semantic Masking
arXiv:2605.25198v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a powerful paradigm for improving language models on reasoning-intensive tasks, but i
Knowledge catalogue
arXiv:2605.25198v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a powerful paradigm for improving language models on reasoning-intensive tasks, but i
arXiv:2208.14882v2 Announce Type: replace-cross Abstract: This paper studies the multimedia problem of temporal sentence grounding (TSG), which aims to accurately determine the specific video segment
arXiv:2605.24763v1 Announce Type: new Abstract: This work presents a high-fidelity computational fluid dynamics (CFD) and data-driven modeling framework for assembly-level flow characterization in a f
arXiv:2605.23922v1 Announce Type: cross Abstract: The EU Artificial Intelligence Act (AIA) establishes a lifecycle governance regime for high-risk AI systems built around ex-ante conformity assessment
arXiv:2509.02113v2 Announce Type: replace-cross Abstract: The advancement of graph-based malware analysis is critically limited by the absence of large-scale datasets that capture the inherent hierarc
arXiv:2605.24635v1 Announce Type: new Abstract: Medical large language models hold promise for reducing healthcare disparities, yet Hindi remains severely underrepresented. While medical LLMs excel in
arXiv:2508.03104v3 Announce Type: replace-cross Abstract: Contrastive learning (CL) has become a dominant paradigm for self-supervised hypergraph learning, enabling effective training without costly l
arXiv:2605.25790v1 Announce Type: new Abstract: The increasing use of drones in human-centric applications highlights the need for designs that can survive collisions and recover rapidly, minimizing r
arXiv:2605.24687v1 Announce Type: cross Abstract: Text-to-Image (T2I) models have made significant strides in visual realism and semantic consistency, yet they often perpetuate and amplify societal bi
Kate Knibbs / Wired: How AI startups like Altur are using chatbots to help automate debt collection; YC incubated six debt collection and settlement startups in the past six years — There's a mad dash
arXiv:2503.07482v2 Announce Type: replace-cross Abstract: Membership Inference Attacks (MIAs) aim to estimate whether a specific data point was used in the training of a given model. Existing state-of
arXiv:2605.24660v1 Announce Type: cross Abstract: Before an LLM agent can use a tool, a retrieval system must decide which candidate tools to show to the agent. How long should that shortlist be? Show
arXiv:2507.19219v2 Announce Type: replace Abstract: Overestimation in evaluating large language models (LLMs) has become an increasing concern. Due to the contamination of public benchmarks or imbalan
arXiv:2605.24351v1 Announce Type: new Abstract: Large language models (LLMs) can support scientific literature synthesis, but remain prone to hallucinated references, uneven coverage, and weakly groun
arXiv:2605.23926v1 Announce Type: new Abstract: Reasoning-capable large language models solve hard problems by emitting long chains of thought, paying heavily in latency, GPU time, and energy. Casual
arXiv:2605.24749v1 Announce Type: cross Abstract: Reward modeling is not only a prediction problem: in KL-regularized policy optimization, the learned reward is exponentiated to define the deployed po
arXiv:2605.25698v1 Announce Type: cross Abstract: High-quality data is scarce in large language model (LLM) training, yet how to schedule its use jointly with training dynamics lacks theoretical guida
arXiv:2605.25414v1 Announce Type: new Abstract: Distribution shift in imitation learning refers to the problem that the agent cannot plan proper actions for a state that has not been visited during th
Most production AI features don't need a frontier model. Here's how capability evals and prompt engineering can help ship a local SLM that matches frontier-model quality with lower latency and cost. T
How to use Grok Build Beginner video: How to install & use Grok Build (made for non-technical SuperGrok and X Premium+ users) I got so many questions from friends, so I made this simple step-by-step g
Over the last 25 years of building Google’s global network, we’ve navigated major architectural eras — from the Internet, to streaming, and the cloud. Today, we are squarely in the midst of a fourth:
arXiv:2605.24229v1 Announce Type: new Abstract: Frontier AI developers now train models against long written behavioral specifications, such as Anthropic's constitution (Anthropic, 2025a) and OpenAI's
Huge one for developers building with AI. Excited to announce @ClementDelangue (CEO @huggingface) is joining DASH 2026 for a fireside chat with @oliveur. Open source changed software, open models are
Huge thanks to @ComfyUI for publishing a deep-dive case study on how we’re building enterprise-grade, scalable AI pipelines. Arron Award Winning work for AI workflow. Read the full breakdown: https://
arXiv:2605.24180v1 Announce Type: cross Abstract: Collaboration is the defining mode of modern science, yet its core mechanism -- feedback -- remains hard to observe, difficult to scale, and unequally
Ivan Mehta / TechCrunch: Human Archive, which trains robots using first-person video from 1,000+ camera-equipped caps worn by Indian home services workers, raised $8.2M from YC and more — In the last
Human in the loop 🤖 Made in @ComfyUI with a laundry list of tools: @ltx_model 2.3 + @thesystms FLW LoRA, @Alibaba_Wan 2.2 I2V + T2V, Florence2, @Meta Sapiens2, @OpenAI GPT Image 2.0, @suno , @AdobeAE
arXiv:2605.24934v1 Announce Type: cross Abstract: Human egocentric video captures rich manipulation demonstrations without any robot hardware, yet transferring these skills to robots remains challengi
arXiv:2605.25685v1 Announce Type: new Abstract: Robust and accurate perception of humans in their 3D scene context is essential for integrating robots into everyday environments. Existing approaches,
arXiv:2508.19113v3 Announce Type: replace Abstract: Large reasoning models (LRMs) combined with retrieval-augmented generation (RAG) have enabled deep research agents capable of multi-step reasoning w
arXiv:2507.02215v2 Announce Type: replace-cross Abstract: Motivated by the need for efficient estimation of conditional expectations, we consider a least-squares function approximation problem with he
arXiv:2603.08072v2 Announce Type: replace Abstract: Forecasting physiological signals can support proactive monitoring and timely clinical intervention by anticipating critical changes in patient stat
arXiv:2605.25421v1 Announce Type: new Abstract: Communication protocol design is a central challenge in large language model-based multi-agent systems. Existing single-channel approaches face an inher
arXiv:2605.24728v1 Announce Type: new Abstract: Foundation models can increasingly describe, reconstruct, and generate 3D objects, assemblies, scenes, and environments, but visually plausible spatial
arXiv:2605.24140v1 Announce Type: new Abstract: Multi-step reasoning remains a central challenge for large language models: single-pass generation is efficient but lacks accuracy; tree-search methods
arXiv:2512.10506v3 Announce Type: replace-cross Abstract: Endmember extraction from hyperspectral images aims to identify the spectral signatures of materials present in a scene. Recent studies have s
arXiv:2605.24528v1 Announce Type: new Abstract: Real world decision-making requires constructing mental models under uncertainty over evidence, over the underlying causal rules, and over the state of
I don't comment on every article that uses out-of-date measures on AI ability, but I felt (probably wrongly) that the article was a response to my viral tweet, so I felt I needed to say something! htt
I feel like I'm eating crazy pills when I read the countless bad takes around how the Vatican would have virtually anointed Anthropic. When if you read the Pope's encyclical it's actually a COMPLETE r
I found this Wired article on AI fact-checking frustrating. It could have been about why we continue to need human fact checkers (talk to people, use judgement, resolve conflict). Instead it is full o
I have been very impressed by @SemiAnalysis_ . I think of myself as a wide ranging systems engineer, looking for value at every level from the chip specs to the user interface, but SA exposes me to ad
I quit ChatGPT for a free, private, and local AI called Ollama - here's why https://www.zdnet.com/article/reasons-to-use-ollama-instead-of-chatgpt/?taid=6a15d1b1cd9b3d000140f77c&utm_campaign=trueAnthe
This post summarizes key financial metrics from SpaceX's hypothetical IPO filing, highlighting concerning metrics including a 700% increase in losses, decelerating revenue growth, and an extremely hig
This post discusses how fundamental data structures in early Unix—specifically three key structs from John Lions' Unix source code commentary—served as the conceptual foundation for the evolution of p
I think folk are underestimating how much of AI models are actually engineering at scale versus breakthrough research. See how @cursor_ai caught up to Anthropic / OpenAI models run at a fraction of th
I wrote a new post on what we need to keep human and what to hand over to AI, with forays into experiments in education, consulting, and the the latest controversy over literary prizes. https://www.on
arXiv:2605.24217v1 Announce Type: new Abstract: As Large Language Models (LLMs) transition from research environments to production deployments, evaluating their performance against strict Service Lev
⚠️⚠️⚠️if you believe this you and think you are safe, you don’t understand how S&P is about to change the rules and what that means. run don’t walk to read my essay “This one weird trick might cost yo
Gary Marcus discusses a potentially overlooked financial risk that could have massive implications for retirement savings, presenting his analysis in an essay format on X. The post appears to emphasiz
I'm going to Applied AI Conference by @techeurope_ in Berlin on Thursday! I'll be having my presentation on the side stage along with other great speakers from @OpenAI, @GoogleDeepMind, @arizeai, @str
- image or video editing? write scripts - finances, tax work, etc? put in PDFs, write scripts, output HTML - medical advice? put in PDFs + data, output HTML - filling out paperwork? write scripts - cr
arXiv:2605.24502v1 Announce Type: cross Abstract: We introduce a physics-inspired continuous relaxation framework that yields substantially improved solutions for NP-hard combinatorial optimization pr
arXiv:2512.23956v2 Announce Type: replace-cross Abstract: Flow Matching (FM) has emerged as a powerful paradigm for continuous normalizing flows, yet standard FM implicitly performs an unweighted L^2
arXiv:2605.25770v1 Announce Type: new Abstract: Robotic systems with redundant degrees of freedom can achieve the same task outcome using multiple configurations, resulting in solution sets that form
Import AI 458 discusses perspectives on AI's future trajectory and potential long-term scenarios, likely including analysis of singularity concepts and their implications. The newsletter entry examine
arXiv:2603.05691v2 Announce Type: replace Abstract: It is increasingly common in machine learning to use learned models to label data and then employ such data to train more capable models. The phenom
arXiv:2605.24009v1 Announce Type: cross Abstract: Convective available potential energy (CAPE) is an important variable for forecasting severe weather and understanding deep convection and precipitati
arXiv:2605.24247v1 Announce Type: cross Abstract: Many automated labeling pipelines classify inputs into categories defined by a written specification, content moderation being a prominent use case. S
arXiv:2605.23924v1 Announce Type: new Abstract: Segment-level disclosures are a central component of financial reporting, providing insight into firms' internal organization and the allocation of econ
In fact, I would be sympathetic to the conclusion that we probably need even more fact checkers, and much more of their time can be freed up to do complex and interesting work by using AI for first-pa