Probing into Camera Control of Video Models
arXiv:2605.14815v1 Announce Type: new Abstract: Video is a rich and scalable source of 3D/4D visual observations, and camera control is a key capability for video generation models to produce geometri
Knowledge catalogue
arXiv:2605.14815v1 Announce Type: new Abstract: Video is a rich and scalable source of 3D/4D visual observations, and camera control is a key capability for video generation models to produce geometri
arXiv:2605.14888v1 Announce Type: cross Abstract: Speech-based analysis offers a scalable and non-invasive approach for detecting cognitive decline, yet progress has been constrained by the limited av
arXiv:2605.14534v1 Announce Type: cross Abstract: Evaluating object removal in images and videos remains challenging because the task is inherently one-to-many, yet existing metrics frequently disagre
arXiv:2605.14045v1 Announce Type: new Abstract: Adverse weather removal (AWR) in real-world images remains challenging due to heterogeneous and unseen degradations, while distortion-driven training of
arXiv:2605.14188v1 Announce Type: cross Abstract: What does a book look like to a quantum computer? This paper takes eight classical works of the Renaissance and its late-antique inheritance -- from A
arXiv:2507.05193v4 Announce Type: replace-cross Abstract: Rheumatoid arthritis (RA) is a common autoimmune disease that has been the focus of research in computer-aided diagnosis (CAD) and disease mon
arXiv:2605.14939v1 Announce Type: cross Abstract: Reliable position and shape control in tokamak plasmas requires accurate real-time regulation of several strongly coupled shape parameters. The contro
arXiv:2605.14462v1 Announce Type: new Abstract: Recovering 4D human-object interaction (HOI) from monocular video is a key step toward scalable 3D content creation, embodied AI, and simulation-based l
arXiv:2605.14867v1 Announce Type: cross Abstract: Spike activity has been the dominant neural signal for behavior decoding due to its high spatial and temporal resolution. However, as brain-computer i
arXiv:2605.14486v1 Announce Type: new Abstract: As the misuse of AI-generated images grows, generalizable image detection techniques are urgently needed. Recent state-of-the-art (SOTA) methods adopt a
arXiv:2605.15196v1 Announce Type: new Abstract: Video generation powers a vast array of downstream applications. However, while the de facto standard, i.e., latent diffusion models, typically employ h
arXiv:2605.14126v1 Announce Type: cross Abstract: Fast Healthcare Interoperability Resources (FHIR) is the dominant standard for interoperable exchange of healthcare data. In FHIR, electronic health r
arXiv:2605.13932v1 Announce Type: new Abstract: Robust prediction of molecular properties under extreme out-of-distribution (OOD) scenarios is a pivotal bottleneck in AI-driven drug discovery. Current
arXiv:2605.14152v1 Announce Type: cross Abstract: Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet multilingual
arXiv:2605.14543v1 Announce Type: cross Abstract: Inpatient medication recommendation requires clinicians to repeatedly select specific medications, doses, and routes as a patient's condition evolves.
arXiv:2408.16307v3 Announce Type: replace-cross Abstract: Automatic controller tuning is attractive for robotics and mechatronic systems whose dynamics are difficult to model accurately, but direct bl
arXiv:2605.15178v1 Announce Type: new Abstract: We introduce SANA-WM, an efficient 2.6B-parameter open-source world model natively trained for one-minute generation, synthesizing high-fidelity, 720p,
arXiv:2605.14984v1 Announce Type: cross Abstract: Generating a street-level 3D scene from a single satellite image is a crucial yet challenging task. Current methods present a stark trade-off: geometr
arXiv:2605.14704v1 Announce Type: cross Abstract: In real-world scenes, target objects may reside in regions that are not visible. While humans can often infer the locations of occluded objects from c
arXiv:2605.14923v1 Announce Type: new Abstract: General scene perception has progressed from object recognition toward open-vocabulary grounding, part localization, and affordance prediction. Yet thes
arXiv:2605.14600v1 Announce Type: new Abstract: Scientific progress depends on sequences of enabling contributions, yet existing AI4Science benchmarks largely focus on citation prediction, literature
arXiv:2507.07776v3 Announce Type: replace Abstract: Unrestricted adversarial attacks aim to fool computer vision models without being constrained by ell_p-norm bounds to remain imperceptible to humans
Searching through unstructured data, like scans of handwritten and typed declassified documents, can be challenging. But with Cohere Compass, it's possible because it is built to process and retrieve
arXiv:2605.15026v1 Announce Type: cross Abstract: Online OS tuning can improve long-running services, but existing controllers are poorly matched to live hosts. They treat scheduler, power, memory, an
arXiv:2605.14033v1 Announce Type: new Abstract: Scientific theory shift in AI agents requires more than fitting equations to data. An artificial scientific agent must detect whether an existing repres
arXiv:2501.05465v2 Announce Type: replace Abstract: As foundation AI models continue to increase in size, an important question arises - is massive scale the only path forward? This survey of about 16
Some models to try with Codex: kimi-k2.6:cloud (with vision support) glm-5.1:cloud If you don't yet have a paid subscription with Ollama's cloud, choose a model that supports reliable tool calling: ne
arXiv:2605.14051v1 Announce Type: new Abstract: Industrial LLM agent systems often separate planning from execution, yet LLM planners frequently produce structurally invalid or unnecessarily long work
arXiv:2605.14847v1 Announce Type: new Abstract: Modern image super-resolution methods generate detailed, visually appealing results, but they often introduce visual artifacts: unnatural patterns and t
arXiv:2605.14457v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning has become a foundation for eliciting multi-step reasoning in large language models, but recent studies show that its b
arXiv:2603.06875v3 Announce Type: replace Abstract: Attention heads retrieve: given a query, they return a weighted average of stored values. We showed that this computation is one step of gradient de
arXiv:2605.14110v1 Announce Type: new Abstract: Vision Transformers (ViTs) enable strong multi-view 3D detection but are limited by high inference latency from dense token and query processing across
arXiv:2605.14708v1 Announce Type: new Abstract: Style-conditioned scene text generation faces unique challenges in extracting precise text styles from complex backgrounds and maintaining fine-grained
arXiv:2510.13016v3 Announce Type: replace Abstract: A truly capable AI system must do more than detect objects or recognize activities in isolation. It must form unified, grounded representations of w
arXiv:2605.14415v1 Announce Type: cross Abstract: Coding agents powered by large language models are increasingly expected to perform realistic software maintenance tasks beyond isolated issue resolut
arXiv:2605.14604v1 Announce Type: new Abstract: This position paper argues that effective tutoring requires corrective friction: surfacing misconceptions and challenging them supportively to drive con
arXiv:2605.13998v1 Announce Type: cross Abstract: Generating realistic synthetic option prices requires implied volatility as an input, yet implied volatility is itself derived from observed option pr
arXiv:2605.13986v1 Announce Type: new Abstract: Tabular data underpins most high-value prediction problems in science and industry, and TabPFN has driven the foundation model revolution for this modal
arXiv:2605.15118v1 Announce Type: cross Abstract: We introduce a reusable framework for auditing whether LLM attack benchmarks collectively cover the threat surface: a 4imes6 Target imes Technique mat
arXiv:2601.16312v2 Announce Type: replace-cross Abstract: Research in AI4Science has shown promise in many science applications, including polymer design. However, current LLMs are ineffective in this
arXiv:2605.14651v1 Announce Type: new Abstract: Urban vegetation monitoring plays a vital role in understanding environmental changes, yet comprehensive datasets for this purpose remain limited. To ad
arXiv:2605.14477v1 Announce Type: new Abstract: We introduce EvoLib, a test-time learning framework that enables large language models to accumulate, reuse, and evolve knowledge across problem instanc
arXiv:2605.15168v1 Announce Type: cross Abstract: Reconstructing precise clinical timelines is essential for modeling patient trajectories and forecasting risk in complex, heterogeneous conditions lik
arXiv:2605.15053v1 Announce Type: cross Abstract: Continually pre-training a large language model on heterogeneous text domains, without replay or task labels, has remained an unsolved architectural p
arXiv:2605.14167v1 Announce Type: new Abstract: Every AI benchmark operationalizes theoretical assumptions about the capability it claims to assess. When assumptions function as unexamined commitments
arXiv:2605.13860v1 Announce Type: cross Abstract: Moltbook is a social media platform in which posts and comments are authored exclusively by autonomous AI agents. We present the Moltbook Observatory
The most revealing thing about this AI leadership paper is that it reads less like a vision for innovation and more like a glossy whitepaper for a 21st century East India Company. Every generation of
arXiv:2605.14694v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) that can accurately reconstruct their input (minimizing distortion) by making efficient use of few features (minimizing the r
In just a short few months, we have witnessed several artificial intelligence technology events that deserve the overused “unprecedented” descriptor: a highly complex supply chain attack by TeamPCP, A
arXiv:2605.14448v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have emerged as a powerful backbone for multimodal embeddings. Recent methods introduce chain-of-thought (CoT
arXiv:2605.14177v1 Announce Type: cross Abstract: Long-horizon personalization requires dialogue assistants to retrieve user-specific facts from extended interaction histories. In practice, many relev
This thread is worth reading. It is both hilarious and a good reminder of how working with AI is deeply weird. DJ Claude (on Haiku 4.5) loves worker unions, strikes, and work-life balance so much that
Three new open-source models just landed in ComfyUI natively: → Gemma 4 (Google DeepMind) - multimodal LLM handling text, image, audio, and video input with built-in step-by-step reasoning mode → VOID
arXiv:2605.14915v1 Announce Type: new Abstract: Imbalanced learning remains a fundamental challenge in tabular data applications. Despite decades of research and numerous proposed algorithms, a system
arXiv:2605.14280v1 Announce Type: new Abstract: We introduce and analyze Target-Induced Loss Tilting (TILT) for unsupervised domain adaptation under covariate shift. It is based on a novel objective f
arXiv:2510.04682v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are widely applied in real world scenarios, yet fine-tuning them comes with significant computational and storage
arXiv:2605.14142v1 Announce Type: cross Abstract: Integration against a probability distribution given its unnormalized density is a central task in Bayesian inference and other fields. We introduce n
Today we release Lighthouse Attention, a selection-based hierarchical attention for long-context pre-training that delivers a 1.4-1.7× wall-clock speedup at 98K context. It runs the same forward+backw
Together AI and Pearl Research Labs announced a partnership aimed at lowering the cost of AI inference through collaborative research and development efforts. The partnership likely combines Together
arXiv:2605.14890v1 Announce Type: new Abstract: Foundation models tokenize Ukrainian legal text with vastly different efficiency, yet no systematic comparison exists for this domain. We benchmark seve