CollaFuse: Collaborative Diffusion Models
arXiv:2406.14429v3 Announce Type: replace-cross Abstract: In the landscape of generative artificial intelligence, diffusion-based models have emerged as a promising method for generating synthetic ima
Knowledge catalogue
arXiv:2406.14429v3 Announce Type: replace-cross Abstract: In the landscape of generative artificial intelligence, diffusion-based models have emerged as a promising method for generating synthetic ima
arXiv:2510.00233v2 Announce Type: replace Abstract: Scientific machine learning has enabled the extraction of physical insights and data-driven modeling of high-dimensional spatiotemporal data, yet ac
arXiv:2510.18900v2 Announce Type: replace-cross Abstract: Accurate prediction of atomistic, thermodynamic, and kinetic properties from molecular structures underpins materials innovation. Existing com
Introducing nanowhale 🐳! A tiny DeepSeek model fully pretrained by an agent. Inspired by @karpathy's nanochat, we gave ml-intern the task of training a tiny MoE with all the architectural advancements
arXiv:2605.00583v1 Announce Type: new Abstract: The visual modality of vision-language models (VLMs) is an underexplored attack surface for bypassing safety alignment. We introduce four jailbreak atta
arXiv:2511.08156v2 Announce Type: replace Abstract: Land Use and Land Cover (LULC) mapping is a fundamental task in Earth Observation (EO). However, current LULC models are typically developed for a s
arXiv:2505.22003v2 Announce Type: replace Abstract: In India, access to legal assistance for the general public has been observed to have a critical gap, as many citizens are not able to take full adv
arXiv:2509.24496v3 Announce Type: replace Abstract: The explosive growth of large language models (LLMs) has created a vast but opaque landscape: millions of models exist, yet their evolutionary relat
🤯 Ollama now supports Claude Desktop via Claude’s built-in third party inference. ollama launch claude-desktop This allows all models from Ollama's Cloud to be used across Claude Cowork and Claude Cod
arXiv:2605.00323v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) often suffer from hallucinations, generating descriptions that include visual details absent from the input image.
I'm obsessed with running local LLMs. Been working with an engineer to build product(s) that are 100% local. A new model that came out recently instantly improved the quality of our product. We live i
This post likely discusses open source AI and language models, potentially covering developments in model accessibility, licensing, or comparisons between proprietary and open source approaches. Given
This article examines how value in the AI industry is shifting away from traditional chip manufacturers toward model development labs and companies that build large language models. It likely analyzes
arXiv:2604.27253v1 Announce Type: new Abstract: Recent advances in multimodal large language models (LLMs) have revolutionized web agents that can automate complex tasks on websites. However, their ac
arXiv:2409.16808v1 Announce Type: cross Abstract: Modern applications, such as autonomous vehicles, require deploying deep learning algorithms on resource-constrained edge devices for real-time image
arXiv:2604.27251v1 Announce Type: cross Abstract: Large Language Models (LLMs) are known to acquire reasoning capabilities through shared inference patterns in pre-training data, which are further eli
arXiv:2604.27929v1 Announce Type: new Abstract: With the widespread adoption of large language models (LLMs), understanding their personality representation mechanisms has become critical. As a novel
arXiv:2503.01611v3 Announce Type: replace Abstract: Instruction following is a critical ability for Large Language Models to perform downstream tasks. The standard approach to instruction tuning has r
arXiv:2604.27607v1 Announce Type: new Abstract: We present JaiTTS-v1.0, a state-of-the-art Thai voice cloning text-to-speech model built through continual training on a large Thai-centric speech corpu
arXiv:2604.27032v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become an integral part of many real-world workflows. However, LLMs consume a lot of energy, which becomes a large c
arXiv:2601.22228v2 Announce Type: replace-cross Abstract: We study whether vision-language models (VLMs) can solve relative camera pose estimation (RCPE) from image pairs, a direct test of multi-view
A Reddit discussion in the r/ollama community seeking recommendations for RAG (Retrieval-Augmented Generation) models, likely addressing model selection for local LLM-based document retrieval systems.
arXiv:2511.17176v2 Announce Type: replace-cross Abstract: Accurate prediction of extreme weather events remains a major challenge for artificial intelligence-based weather prediction systems. While de
One week since the launch of GPT-5.5, and it’s already our strongest model launch yet. API revenue is growing more than 2x faster than any prior release, while Codex doubled revenue in under seven day
arXiv:2506.02671v3 Announce Type: replace Abstract: Deep neural networks often suffer performance degradation upon deployment due to distribution shifts. Continual Test-Time Adaptation (CTTA) aims to
arXiv:2510.17196v3 Announce Type: replace-cross Abstract: Effectively processing long contexts is a critical challenge for language models. While standard Transformers are limited by quadratic complex
When training Grok 4.3, we spoke directly with devs and businesses to understand what they actually needed: a model that’s fast, affordable, and great at tool calling. The result is a daily driver tha
5/5 Maestro automatically searches across this space (model ensembles, scaling strategies, execution policies), maps the Pareto frontier, and surfaces the full accuracy–cost–latency tradeoff surface.
arXiv:2604.26348v1 Announce Type: cross Abstract: Diffusion models have achieved remarkable success in image generation, yet their training is predominantly driven by full-reference objectives that en
arXiv:2604.26404v1 Announce Type: new Abstract: Industrial object detection systems typically rely on large annotated datasets, which are expensive to collect and challenging to maintain in industrial
arXiv:2604.26419v1 Announce Type: cross Abstract: Large Vision-Language Models (VLMs) have achieved remarkable multimodal performance yet remain prone to factual hallucinations, particularly in long-t
In a federal courtroom in California on Thursday, Elon Musk testified that his own AI startup, xAI, has used OpenAI's models to improve its own. The matter at question is model distillation, a common
arXiv:2604.26182v1 Announce Type: cross Abstract: World models of embodied agents predict future observations conditioned on an action taken by the agent. For complex embodiments, action spaces are hi
Mistral AI made the TIME100 Most Influential Companies list for 2026 — and the top 10 for AI. Why we're proud: customers run frontier models in production on their own terms, on their own infrastructu
PEOPLE ARE NOW RUNNING CLAUDE CODE WITH LOCAL AI MODELS TO AVOID API COSTS. By connecting tools like Ollama and Gemma 4, developers can build apps locally with unlimited usage and no monthly billing.
arXiv:2604.26078v1 Announce Type: new Abstract: Photoplethysmography (PPG) is increasingly used in wearable affective computing due to its low cost and ease of integration into consumer devices. Recen
arXiv:2604.26508v1 Announce Type: cross Abstract: Deploying Vision-Language Models (VLMs) on edge devices remains challenging due to their substantial computational and memory demands, which exceed th
arXiv:2512.23726v2 Announce Type: replace-cross Abstract: The 3D fast silent multi-parametric mapping sequence with zero echo time (MuPa-ZTE) is a novel quantitative MRI (qMRI) acquisition that enable
arXiv:2604.24966v1 Announce Type: cross Abstract: Frontier AI companies first deploy their most advanced models internally, for weeks or months of safety testing, evaluation, and iteration, before a p
Train a TensorFlow object detection model – then deploy it on a robot 🤖 Iulia Feroli (@iuliaferoli) shows how to turn a notebook into a real-time object detection app. This tutorial works for any proj
arXiv:2604.26694v1 Announce Type: cross Abstract: We propose X-WAM, a Unified 4D World Model that unifies real-time robotic action execution and high-fidelity 4D world synthesis (video + 3D reconstruc
arXiv:2603.07080v3 Announce Type: replace-cross Abstract: Vision-and-Language Navigation (VLN) increasingly relies on large vision-language models, but their inference cost conflicts with real-time de
we're starting rollout of GPT-5.5-Cyber, a frontier cybersecurity model, to critical cyber defenders in the next few days. we will work with the entire ecosystem and the government to figure out trust
// When to Retrieve During Reasoning // Pay attention to this one, AI devs. (bookmark it) Most RAG systems retrieve once, before the model starts reasoning. Large reasoning models like o1 and R1 don't
arXiv:2604.25903v1 Announce Type: cross Abstract: The accelerating adoption of Large Language Models (LLMs) in software engineering (SE) has brought with it a silent crisis: unsustainable computationa
arXiv:2602.01785v2 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved remarkable success in source code understanding, yet as software systems grow in scale, computational eff
arXiv:2604.24959v1 Announce Type: new Abstract: Learning matrix-valued distributions from high-dimensional and possibly incomplete training data is challenging: ambient-space generative modeling is co
arXiv:2604.25853v1 Announce Type: new Abstract: Traditional loss functions, including cross-entropy, contrastive, triplet, and su pervised contrastive losses, used for fine-tuning pre-trained language
IBM Granite just released two multilingual embedding models with 97M and 311M parameters 🤏🏻 ModernBERT-based, 200+ languages, 32K context, and built for retrieval, search, similarity, and code. And...
arXiv:2604.24765v1 Announce Type: cross Abstract: P300-based brain-computer interfaces (BCIs) are widely used for communication, but population heterogeneity may alter the neural patterns available fo
IBM Granite 4.1 is a family of foundation models released by IBM Research designed for enterprise AI applications. The models likely offer improvements in performance, efficiency, and capabilities com
Its a bit frustrating, because Gemini 3.1 Pro is an excellent model and can deliver really good results. But here is GPT-5.5 Pro for comparison. Sadly, it took this assignment very seriously and ethic
arXiv:2604.25765v1 Announce Type: new Abstract: The quality of training data is critical to the performance of machine learning models. In this paper, the Error Sensitivity Profile (ESP) is proposed.
arXiv:2604.25674v1 Announce Type: new Abstract: Modeling the emergence of human-like lexicons in computational systems has advanced through the use of interacting neural agents, which simulate both le
arXiv:2602.17697v2 Announce Type: replace Abstract: Large Language Models (LLMs) are being increasingly used across a wide range of tasks. However, their substantial computational demands raise concer
arXiv:2604.25154v1 Announce Type: new Abstract: Tabular Foundation Models (TFMs) achieve state-of-the-art zero-shot accuracy on small tabular datasets by meta-learning over synthetic data-generating p
arXiv:2604.24175v1 Announce Type: cross Abstract: Large language models have demonstrated strong reasoning capabilities in general knowledge question answering. However, their ability to handle tempor
arXiv:2508.10599v4 Announce Type: replace Abstract: Activation steering offers a promising approach to controlling the behavior of Large Language Models by directly manipulating their internal activat
arXiv:2512.21110v3 Announce Type: replace Abstract: Current Large Language Models (LLMs) safety approaches focus on explicitly harmful content while overlooking a critical vulnerability: the inability
arXiv:2604.24273v1 Announce Type: new Abstract: The deployment of intelligent reinforcement learning (RL) agents on resource-constrained edge devices remains a fundamental challenge due to the substan