CompanyAnthropic8 recent entries10 Aug 2026Oh no, we aren’t going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, bec…Oh no, we aren’t going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, because our experiments (with slightly older models) found it d→10 Aug 2026Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference ElicitationarXiv:2608.06955v1 Announce Type: new Abstract: Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs sys
CompanyOpenAI8 recent entries11 Aug 2026Stealing Reasoning Traces from Proprietary LLM APIsarXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim→11 Aug 2026Source: Trajectory, founded by ex-DeepMind, Apple, OpenAI, and Meta staffers to build continual learning models, raised 40M led by Sequoia at a 300M valuation (Stephanie Palazzolo/The Information)Stephanie Palazzolo / The Information: Source: Trajectory, founded by ex-DeepMind, Apple, OpenAI, and Meta staffers to build continual learning models, raised 40M led by Sequoia at a 300M valuation — →11 Aug 2026OpenAI just launched a cybersecurity model that answers 95% of advanced threat queries. And Meta put a frontier model on your laptop. Same day.Something happened today that I think most people are going to miss because there are two separate stories and neither one is getting the full picture. OpenAI expanded Daybreak. If you haven't heard o→11 Aug 2026Introducing Unsloth Desktop appHi LocalLlama, we're super excited to release Unsloth Desktop today! 🦥 It's the first desktop app that enables you to run and train models locally. Open-source. Available on Mac, Windows, and Linux Su→11 Aug 2026Epic talk: the cheat code for how to build your own in-house lab, featuring @gabepereyra of @harveyEpic talk: the cheat code for how to build your own in-house lab, featuring @gabepereyra of @harvey Want world class research capabilities, but don’t have the resources of a big lab? At our recent Sov→11 Aug 2026AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research GroupsarXiv:2608.08883v1 Announce Type: new Abstract: Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific workflows. H→12 Aug 2026Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational EngagementarXiv:2608.10672v1 Announce Type: cross Abstract: Social interaction has become one of the most common uses of LLMs, yet research on emotional bonds with AI has focused largely on how users experience→12 Aug 2026Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show ope…Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show open weights winning over time, but work submitted to Pangram i
CompanyGoogle8 recent entries11 Aug 2026Towards Expert-level Medical AI for Real-time Video ConsultationsarXiv:2608.09861v1 Announce Type: new Abstract: Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through→11 Aug 2026Source: Trajectory, founded by ex-DeepMind, Apple, OpenAI, and Meta staffers to build continual learning models, raised 40M led by Sequoia at a 300M valuation (Stephanie Palazzolo/The Information)Stephanie Palazzolo / The Information: Source: Trajectory, founded by ex-DeepMind, Apple, OpenAI, and Meta staffers to build continual learning models, raised 40M led by Sequoia at a 300M valuation — →11 Aug 2026Researchers find that feeding a frontier model's encrypted reasoning traces to a weaker model from the same provider can make it output the traces in plaintext (Will Knight/Wired)Will Knight / Wired: Researchers find that feeding a frontier model's encrypted reasoning traces to a weaker model from the same provider can make it output the traces in plaintext — Researchers devis→11 Aug 2026Looker’s semantic layer governs Gemini Enterprise data for user trustFor organizations deploying AI agents at scale, there’s often a critical divide between structured and unstructured data. While large language models (LLMs) excel at parsing text documents, emails, an→11 Aug 2026Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection PressurearXiv:2608.08722v1 Announce Type: cross Abstract: Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. We document this concretely →12 Aug 2026Reference-Free Post-Training of Open Large Language Models for Multilingual Machine TranslationarXiv:2608.10812v1 Announce Type: cross Abstract: We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiL→12 Aug 2026Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI AgentsarXiv:2608.09944v1 Announce Type: cross Abstract: Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure be→12 Aug 2026Google debuts SL2T, an AI model that’s designed to understand sign languageGoogle DeepMind said today it wants to bring the artificial intelligence revolution to the estimated 70 million people across the world who are either deaf or hard of hearing with the launch of sign-l
CompanyMeta8 recent entries11 Aug 2026iLTM: Integrated Large Tabular ModelarXiv:2511.15941v2 Announce Type: replace-cross Abstract: Tabular data underpins decisions across science, industry, and public services. Despite rapid progress, advances in deep learning have not ful→11 Aug 2026Communication-efficient distributed hazard difference estimation for heterogeneous multi-site survival dataarXiv:2601.14609v2 Announce Type: replace-cross Abstract: Multi-site collaboration can power survival models that no single hospital could fit alone, but privacy rules and protected computing environm→11 Aug 2026Automated Generation of Complexity-Validated Decision Scenarios Using Large Language ModelsarXiv:2608.08822v1 Announce Type: new Abstract: Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and b→11 Aug 2026A New Approach to Characterising Optimisation Problems Using Programmatic Representation and Complexity MeasuresarXiv:2608.08898v1 Announce Type: cross Abstract: Characterising optimisation problem instances is a fundamental part of understanding the behaviour and performance of different algorithms as well as →12 Aug 2026Physics-informed Diffusion Generative Model for Time-Series Data Synthesis in Dynamic SystemsarXiv:2608.10941v1 Announce Type: new Abstract: Industrial time-series signals, such as turbine temperature and rotational speed in aero-engines, are essential for monitoring the health and operationa→12 Aug 2026New Muse-Glimmer-30B SoTA Quants - hopefully a new lineup :)Hey Folks, I've been making quants for a while - recently I took a short break to get into hardcore research (submitted my first EMNLP paper during it!). Along the way, I built up a little arsenal of →12 Aug 2026Muse Glimmer is live on Fireworks. The new open-weight model from Meta Superintelligence Labs is a 30B dense model built for always-on agent…Muse Glimmer is live on Fireworks. The new open-weight model from Meta Superintelligence Labs is a 30B dense model built for always-on agents that reason across many sequential tool calls and can reco→12 Aug 2026Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent SystemsarXiv:2608.09964v1 Announce Type: cross Abstract: The Brazilian Conference on Intelligent Systems (BRACIS) is the main national venue for Artificial Intelligence research in Brazil, hosted by the Braz
CompanyMistral8 recent entries31 Jul 2026What’s new in AI infrastructure and orchestration this monthAt Google, AI is a soup-to-nuts endeavor. Obviously, we make leading AI models like Gemini and Nano Banana. We incorporate AI into the tools you use every day (think Gmail, BigQuery, AlloyDB, Google C→31 Jul 2026Sympathetic Framing: Evaluating AI Alignment across Sociodemographic GroupsarXiv:2607.27232v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs→31 Jul 2026STEREODISCO: Discovering Stereotypicality in LLMsarXiv:2607.27824v1 Announce Type: cross Abstract: LLMs encode, convey, and perpetuate stereotypes. Prior computational research focuses on a small set of semantic axes investigated in social psycholog→5 Aug 2026Prime Agent - a new coding harness surpassing Codex/CC/PIPrime Agent is an open-source coding and research agent for general and long-running work. A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficien→7 Aug 2026PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMsarXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden →7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr→10 Aug 2026Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference ElicitationarXiv:2608.06955v1 Announce Type: new Abstract: Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs sys→11 Aug 2026Who Verifies the Benchmark? Decentralizing Trust in Large Language Model EvaluationarXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha
CompanyxAI8 recent entries7 Aug 2026upgraded my stack, and i can now work on almost anything from anywhere hands free: - talk to chief of staff (via remote codex voice or text)…upgraded my stack, and i can now work on almost anything from anywhere hands free: - talk to chief of staff (via remote codex voice or text) - chief assigns tasks to managers of various projects - man→7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr→10 Aug 2026Evaluating XAI Support From A Hierarchical Reinforcement Learning Policy in Human-Agent CollaborationarXiv:2608.06381v1 Announce Type: cross Abstract: Explainable AI (XAI) has shown promise for human-agent collaboration, yet results rely on hand-crafted policies in custom environments, limiting gener→12 Aug 2026The Epistemic Politics of AI AnthropomorphismarXiv:2608.00961v2 Announce Type: replace-cross Abstract: AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction. Users who engage in sustained →12 Aug 2026Rule of Thumb: Explaining Artificial Intelligence Systems using Partial InformationarXiv:2608.10766v1 Announce Type: new Abstract: Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. We propose ''Rul→12 Aug 2026Entropy-Centric Explainable AI for Remote Sensing Image SegmentationarXiv:2608.11064v1 Announce Type: cross Abstract: Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many concerns arise regarding the decisio→12 Aug 2026Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human UnderstandingarXiv:2603.25251v2 Announce Type: replace-cross Abstract: Explainable AI (XAI) methods are commonly evaluated using functional correctness metrics, sometimes termed faithfulness or fidelity, which est→12 Aug 2026Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and ReliancearXiv:2608.10434v1 Announce Type: new Abstract: Machine learning-based Intrusion Detection Systems (IDS) have demonstrated superior performance in securing Unmanned Aerial Vehicle (UAV) networks. Howe
CompanyDeepSeek8 recent entries9 Aug 2026endless-frontier/BigBang-v1 - qwen 3.5 finetunestable bench https://huggingface.co/bartowski/endless-frontier_BigBang-v1-GGUF I'm downloading this model only because Bartowski converted it to .gguf, so it might be interesting. Doubts : The headline→10 Aug 2026WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic GraderarXiv:2608.06474v1 Announce Type: new Abstract: Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central approa→10 Aug 2026Fisher-R1: Training LLM Agents for Reliable Hypothesis TestingarXiv:2608.07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate t→11 Aug 2026Who Verifies the Benchmark? Decentralizing Trust in Large Language Model EvaluationarXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha→11 Aug 2026Other active promotions: - Free models: Solar Pro 4 (1 week), Hy3, Step 3.7 Flash, Laguna S and XS - 90% off DeepSeek V4 Flash for ~2 more d…Nous Research has extended its 20 % discount on all models—including high‑end frontier options—throughout the Nous Portal for an additional two weeks (until the end of April). Free model trials such a→11 Aug 2026Automated Generation of Complexity-Validated Decision Scenarios Using Large Language ModelsarXiv:2608.08822v1 Announce Type: new Abstract: Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and b→12 Aug 2026Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show ope…Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show open weights winning over time, but work submitted to Pangram i→12 Aug 2026Idea for a deepseek-v4-flash-0731 backed automated research workflow to be leveraged via qwen3.6/3.8 27b for difficult tasks that require highly technical, not easy to find information.Sometimes you have tasks that are outside of your expertise and the idea is this workflow automation could be leveraged to manage to have local AI figure it out using research from his workflow gather
CompanyNVIDIA8 recent entries6 Aug 2026Advancing brain tumor research with privacy-first AIThe intersection of medicine and AI has led to remarkable innovations. However, developers now face the thorny challenge of building robust medical AI tools that have been tested and evaluated on dive→9 Aug 2026I Turned My Underused Gaming Laptop Into a Local AI WorkstationTL;DR: I am building a Windows-first local AI setup for people who want to try local LLMs without spending days choosing models, setting up Ollama, Docker, WSL, Open WebUI, agents, and tool permission→10 Aug 2026Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflowsHi r/LocalLLaMA 👋 Today we’re excited to release Muse Glimmer, a 30B open-weight model built specifically for local agent workflows. We’re releasing the weights to the community under a permissive Apa→10 Aug 2026Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows. Muse Glimmer delivers strong pe…Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows. Muse Glimmer delivers strong performance on key agentic use cases and benchmarks compared w→10 Aug 2026Dual-Node NVIDIA DGX Spark over Tailscale: A Remote-Access Testbed for Distributed LLM Training and Cyber-Threat-Intelligence Fine-TuningarXiv:2608.07226v1 Announce Type: cross Abstract: Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class →11 Aug 2026Who Verifies the Benchmark? Decentralizing Trust in Large Language Model EvaluationarXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha→11 Aug 2026Introducing Unsloth Desktop appHi LocalLlama, we're super excited to release Unsloth Desktop today! 🦥 It's the first desktop app that enables you to run and train models locally. Open-source. Available on Mac, Windows, and Linux Su→11 Aug 2026EsaacSim: A Multimodal Event Camera Add-on for NVIDIA Isaac SimarXiv:2608.08522v1 Announce Type: new Abstract: Event-based vision is becoming an increasingly important sensing paradigm for robotics, yet its adoption remains limited by sensor availability and the