R3D: Revisiting 3D Policy Learning
arXiv:2604.15281v1 Announce Type: new Abstract: 3D policy learning promises superior generalization and cross-embodiment transfer, but progress has been hindered by training instabilities and severe o
Knowledge catalogue
arXiv:2604.15281v1 Announce Type: new Abstract: 3D policy learning promises superior generalization and cross-embodiment transfer, but progress has been hindered by training instabilities and severe o
arXiv:2604.13757v1 Announce Type: new Abstract: The next generation of autonomous AI systems will be constrained not only by model capability, but by how intelligence is structured across heterogeneou
arXiv:2604.14656v1 Announce Type: cross Abstract: Most medical multimodal benchmarks focus on static tasks such as image question answering, report generation, and plain-language rewriting. Patient ed
arXiv:2512.00995v3 Announce Type: replace Abstract: Part-level point cloud segmentation has recently attracted significant attention in 3D computer vision. Nevertheless, existing research is constrain
arXiv:2603.16024v2 Announce Type: replace Abstract: We introduce a speech-guided embodied agent framework for video-guided skull base surgery that dynamically executes perception and image-guidance ta
arXiv:2604.13041v1 Announce Type: cross Abstract: Table Structure Recognition (TSR) requires the logical reasoning ability of large language models (LLMs) to handle complex table layouts, but current
arXiv:2604.14619v1 Announce Type: cross Abstract: In computational paralinguistics, detecting cognitive load and deception from speech signals is a heavily researched domain. Recent efforts have attem
arXiv:2604.15301v1 Announce Type: new Abstract: Many SLT systems quietly assume that brief chunks of signing map directly to spoken-language words. That assumption breaks down because signers often cr
arXiv:2507.15066v5 Announce Type: replace Abstract: Time series anomaly detection (TSAD) has traditionally focused on binary classification and often lacks the fine-grained categorization and explanat
arXiv:2511.07412v2 Announce Type: replace Abstract: Developing embodied AI for intelligent surgical systems requires safe, controllable environments for continual learning and evaluation. However, saf
arXiv:2505.20291v4 Announce Type: replace-cross Abstract: Text-to-image retrieval (T2I retrieval) remains challenging because cross-modal embeddings often behave as bags of concepts, underrepresenting
arXiv:2604.14548v1 Announce Type: cross Abstract: As speech language models (SLMs) transition from personal devices into shared, multi-user environments, their responses must account for far more than
arXiv:2604.15010v1 Announce Type: cross Abstract: When do transformers commit to a decision, and what prevents them from correcting it? We introduce extbf{prolepsis}: a transformer commits early, task
arXiv:2604.14749v1 Announce Type: new Abstract: Large language models still struggle with faithfulness and hallucinations despite their remarkable reasoning abilities. In Knowledge Graph Question Answ
arXiv:2604.14540v1 Announce Type: new Abstract: Detecting slow-moving landslides directly from wrapped Interferometric Synthetic Aperture Radar (InSAR) interferograms is crucial for efficient geohazar
Agent evals are drifting away from production reality. Most benchmarks use clean tasks, well-specified requirements, deterministic metrics, and retrospective curation. Production work is messier, with
Iris Deng / South China Morning Post: Amazon launches its first smart warehouse in Shenzhen, aiming to cut local merchant storage costs by up to 45% as competition with Shein and Temu intensifies — Am
Anthropic says Opus 4.7 hits 80.6% on Document Reasoning — up from 57.1%. But 'reasoning about documents' ≠ 'parsing documents for agents.' We ran it on ParseBench. → Charts: 13.5% → 55.8% (+42.3) — h
arXiv:2604.13795v1 Announce Type: new Abstract: Vision transformers (ViT) have been shown to allow for more flexible feature detection and can outperform convolutional neural network (CNN) when pre-tr
arXiv:2507.16005v2 Announce Type: replace-cross Abstract: Alloy discovery is constrained by vast compositional spaces, competing objectives, and prohibitive experimental costs. Although simulations an
arXiv:2604.13305v1 Announce Type: new Abstract: Reward models (RMs) are inherently non-neutral value functions designed and trained to encode specific objectives, such as human preferences or text-ima
arXiv:2604.13253v1 Announce Type: new Abstract: Adaptive Conformal Inference (ACI) provides distribution-free prediction intervals with asymptotic coverage guarantees for time series under distributio
arXiv:2604.13263v1 Announce Type: new Abstract: Meta-learning offers a principled framework leveraging task-invariant priors from related tasks, with which task-specific models can be fine-tuned on do
arXiv:2604.13304v1 Announce Type: new Abstract: Understanding the internal activations of Vision Transformers (ViTs) is critical for building interpretable and trustworthy models. While Sparse Autoenc
arXiv:2604.13452v1 Announce Type: new Abstract: Long-form visual storytelling requires maintaining continuity across shots, including consistent characters, stable environments, and smooth scene trans
arXiv:2604.13409v1 Announce Type: new Abstract: In clinical practice, the robustness of deep learning models for multimodal brain tumor segmentation is severely compromised by incomplete MRI data. Thi
Claude Opus 4.7 is now available in Windsurf 2.0! Anthropic has clearly optimized Claude Opus 4.7 for sustained reasoning over long runs. Agents stay on track longer without intervention, so engineers
arXiv:2601.15550v3 Announce Type: replace Abstract: Existing cultural commonsense benchmarks treat nations as monolithic, assuming uniform practices within national boundaries. But does cultural commo
arXiv:2604.13761v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) layers have been shown to substantially increase model capacity without a proportional increase in computational cost an
arXiv:2604.13878v1 Announce Type: new Abstract: Driver drowsiness significantly impairs the ability to accurately judge safe braking distances and is estimated to contribute to 10%-20% of road acciden
arXiv:2604.13586v1 Announce Type: new Abstract: Existing multi-view three-dimensional (3D) object detection approaches widely adopt large-scale pre-trained vision transformer (ViT)-based foundation mo
Claude Code users should ensure they are using Opus 4.7 with the latest updates to experience performance improvements or feature enhancements. The post suggests that users not observing expected incr
arXiv:2604.13398v1 Announce Type: new Abstract: While Aspect-based Sentiment Analysis (ABSA) systems have achieved high accuracy in identifying sentiment polarities, they often operate as 'black boxes
arXiv:2507.13942v2 Announce Type: replace Abstract: Forecasting future events is a fundamental capability for general-purpose systems that plan or act across different levels of abstraction. Yet, eval
Gemini now uses user interests and Google Photos to create personalized AI images without requiring long descriptions, allowing users to simply ask for pictures of themselves or family members. This f
arXiv:2604.13954v1 Announce Type: new Abstract: Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions.
In this post, you'll learn why probabilistic AI validation falls short in regulated industries and how Automated Reasoning checks use formal verification to deliver mathematically proven results. You'
Nous Research shared an update or announcement on their official X (Twitter) account, likely relating to their ongoing work in AI model development, fine-tuning, or research releases. Nous Research is
in the grand narrative of Meta x AI, we saw the flop (Llama 4 hurhurhur), and now we’re seeing the turn: - *more* hiring since the soup wars of 2025 - Zuck literally moved in with Alexandr and Nat and
arXiv:2604.13604v1 Announce Type: cross Abstract: Binary stellar evolution simulations are computationally expensive. Stellar population synthesis relies on these detailed evolution models at a fundam
arXiv:2604.13733v1 Announce Type: new Abstract: Reinforcement learning (RL) enables high-frequency, closed-loop control for robotic manipulation, but scaling to long-horizon tasks with sparse or imper
arXiv:2508.06433v4 Announce Type: replace Abstract: Large Language Models (LLMs) based agents excel at diverse tasks, yet they suffer from brittle procedural memory that is manually engineered or enta
This post celebrates Claude Opus 4.7's capabilities within Claude Code, expressing enthusiasm about its performance and describing it as exceptionally powerful. The entry references an Anthropic annou
This paper was accepted at the Workshop on Navigating and Addressing Data Problems for Foundation Models (NADPFM) at ICLR 2026. Principled domain reweighting can substantially improve sample efficienc
Motif-Video-2B is a 2-billion-parameter open-source video generation diffusion transformer released by Motif Technologies in April 2026, capable of both text-to-video and image-to-video generation fro
Thunderbolt is a new open-source AI client from Mozilla-owned MZLA Technologies aimed at enterprises who want to run self-hosted chatbots on their own infrastructure. The platform allows organizations
🎬 Ollama Gemma Day Recap: SGLang at the Ollama Gemma 4 Party in Palo Alto 🍾 Last night, @ollama hosted a packed Gemma Day at the Palo Alto office alongside the @GoogleDeepMind Gemma team. SGLang was i
arXiv:2604.13283v1 Announce Type: cross Abstract: Earth Observation (EO) satellite scheduling (deciding which imaging tasks to perform and when) is a well-studied combinatorial optimization problem. E
Opus 4.7 is in Claude Code today. It's more agentic, more precise, and a lot better at long-running work. It carries context across sessions and handles ambiguity much better. Introducing Claude Opus
arXiv:2604.13175v1 Announce Type: new Abstract: Large language models can be aligned with human preferences through offline reinforcement learning (RL) on small labeled datasets. While single-objectiv
arXiv:2604.13918v1 Announce Type: new Abstract: We present PartNerFace, a part-based neural radiance fields approach, for reconstructing animatable facial avatar from monocular RGB videos. Existing so
arXiv:2604.13153v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has recently enabled highly photorealistic 3D reconstruction from casually captured multi-view images. However, this access
arXiv:2604.13951v1 Announce Type: new Abstract: This study evaluates colorectal risk factors and compares classical models against Quantum Neural Networks (QNNs) for anastomotic leak prediction. Analy
arXiv:2505.19054v2 Announce Type: replace Abstract: Modern learning-based locomotion controllers typically rely on fully trainable deep neural networks with a large number of parameters. This paper st
arXiv:2604.13549v1 Announce Type: new Abstract: The conversion of 2D freehand sketches into 3D models remains a pivotal challenge in computer vision, bridging the gap between human creativity and digi
arXiv:2604.13905v1 Announce Type: new Abstract: We present SparseGen, a novel framework for efficient image-to-3D generation, which exhibits low input-view bias while being significantly faster. Unlik
arXiv:2204.13635v2 Announce Type: replace Abstract: Depth completion involves recovering a dense depth map from a sparse map and an RGB image. Recent approaches focus on utilizing color images as guid
arXiv:2512.15742v2 Announce Type: replace Abstract: Pre-trained Vision Kolmogorov-Arnold Networks (KANs) store a dense B-spline grid on every edge, inflating prediction-head parameter counts by more t
Shocking result on my pelican benchmark this morning, I got a better pelican from a 21GB local Qwen3.6-35B-A3B running on my laptop than I did from the new Opus 4.7! Qwen on the left, Opus on the righ
arXiv:2604.13581v1 Announce Type: new Abstract: Accurately reconstructing human behavior in close-interaction scenarios is crucial for enabling realistic virtual interactions in augmented reality, pre