Wordle 1,812 4/6 ⬛⬛⬛⬛⬛ 🟨⬛🟨⬛🟨 ⬛🟨🟨⬛🟨 🟩🟩🟩🟩🟩
This post shows a completed Wordle game (#1,812) solved in 4 attempts, displaying the colored tile feedback pattern (gray for incorrect letters, yellow for correct letters in wrong positions, green fo
Knowledge catalogue
This post shows a completed Wordle game (#1,812) solved in 4 attempts, displaying the colored tile feedback pattern (gray for incorrect letters, yellow for correct letters in wrong positions, green fo
arXiv:2606.06147v1 Announce Type: new Abstract: End-to-end Vision-Language-Action (VLA) models have shown promise in UAV navigation. However, existing approaches typically rely on historical observati
You can’t let this happen, @DavidSacks, cc @elonmusk. It’s WWWIII but with AI. Nobody wins. ⚠️⚠️ Seismic shift ⚠️⚠️ It’s a good day to be Mistral. Nobody is going to trust an American AI company that
Your margin is my opportunity: AI version… The biggest surprise of 2026 is that the capability gap between the best open-weight/source models and the best closed models has narrowed much faster than t
SpaceX launched 29 additional Starlink satellites, bringing the total constellation to over 10,000 satellites in orbit. This milestone represents a significant expansion of the Starlink mega-constella
arXiv:2606.06077v1 Announce Type: new Abstract: Autonomous underwater vehicle (AUV) launch and recovery (LAR) into the hull of an advancing host platform requires traversal of a complex, three-dimensi
arXiv:2606.05903v1 Announce Type: new Abstract: Performance variations in sensor arrays, caused by intrinsic differences or installation conditions, can lead to inconsistent results during shape sensi
Coco Feng / South China Morning Post: A research team that includes Huawei says it successfully used Huawei's Ascend 910C chips for DeepSeek V4 Pro model's post-training, amid increased US sanctions —
arXiv:2606.05622v1 Announce Type: new Abstract: Planning for real-world problems by language models often involves both world and user constraints, which may not be fully specified upfront and are pro
arXiv:2606.06158v1 Announce Type: new Abstract: Adaptive video tokenisation seeks to dynamically allocate token budgets based on the underlying visual complexity of a sequence. Current continuous-regi
arXiv:2606.05405v1 Announce Type: cross Abstract: Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deploym
// Agents' Last Exam // Agents' Last Exam is a living benchmark of over 1,000 economically valuable tasks, built with 250+ industry experts and mapped to the U.S. federal occupational taxonomy. The ha
arXiv:2602.09574v2 Announce Type: replace Abstract: Tree-search decoding is an effective form of test-time scaling for large language models (LLMs), but real-world deployment often imposes a fixed per
arXiv:2606.05531v1 Announce Type: cross Abstract: Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and ch
arXiv:2606.05725v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed through hosted APIs, making model extraction a practical threat to model ownership and service
OpenAI experienced a technical issue that resulted in some user accounts being incorrectly suspended. The company announced it is working to restore access to affected accounts and address related iss
arXiv:2606.05553v1 Announce Type: new Abstract: Role-playing language agents (RPLAs) should play characters whose values and behavior evolve as the story progresses, not maintain a fixed persona. Exis
Arena AI's agentic benchmark ranks AI models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability. The leaderboard
arXiv:2606.05920v1 Announce Type: cross Abstract: Existing code-generation benchmarks score a single mapping from a complete prompt to a one-shot output. However, real web development is different. Us
At least until (if?) rapid improvement stops, it seems less likely someone is going to catch the Big Three AI Labs. Microsoft and Meta released their models, which were fine, but not frontier. SpaceX
Augment Code Computing Inc., an artificial intelligence agent platform provider, Thursday announced the launch of Cosmos, a service it says is designed to push beyond the era of individual AI coding a
arXiv:2606.05557v1 Announce Type: new Abstract: A situated query like 'where is Lin Wei?' often encodes more than its literal content: the user may also want to know whether Lin Wei is free, in a good
Before the week ends, let's acknowledge one of the most INSANE week ever for open AI, with 25+ notable open-weight drops across every modality: 🧠 LLMs → NVIDIA Nemotron 3 Ultra: 550B hybrid Mamba-MoE,
arXiv:2606.06242v1 Announce Type: new Abstract: Institutional documents contain substantial amounts of operational and analytical information embedded within figures and tables. Current approaches for
arXiv:2606.05924v1 Announce Type: new Abstract: Literary translation poses unique challenges due to the scarcity of high-quality annotated data and the need to balance expression fluency with literary
arXiv:2606.05368v1 Announce Type: new Abstract: Accurate, spatially explicit characterization of tropical forest structure is essential for carbon accounting and ecosystem monitoring, yet most ML pipe
arXiv:2606.06250v1 Announce Type: new Abstract: Continuous-time SLAM provides a principled framework for fusing heterogeneous sensors while estimating smooth trajectories, and is particularly well-sui
arXiv:2606.05915v1 Announce Type: new Abstract: Estimating 2D camera motion is fundamental to computer vision and computational photography. Existing homography-based methods work well for planar scen
arXiv:2606.05804v1 Announce Type: new Abstract: Prompted knowledge cutoff instructs a large language model (LLM) to act as if information beyond a specified cutoff date were unavailable. However, prio
arXiv:2606.06088v1 Announce Type: new Abstract: We present CHALIS (Challenging Language Identification Samples), a new benchmark dataset explicitly designed to address difficult cases in language iden
arXiv:2410.13056v4 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable success across a wide range of language tasks, but their deployment on edge devices remain
arXiv:2504.10823v4 Announce Type: replace Abstract: Navigating dilemmas involving conflicting values is challenging even for humans in high-stakes domains, let alone for AI, yet prior work has been li
arXiv:2606.06219v1 Announce Type: new Abstract: End-to-end autonomous driving models often struggle to balance multi-modal maneuver generation with real-time inference constraints. While diffusion mod
arXiv:2602.23845v2 Announce Type: replace Abstract: Chinese text correction has traditionally focused on spelling and grammar, while factual error correction is usually treated separately. However, in
arXiv:2606.06492v1 Announce Type: cross Abstract: Code language models need repository-level context to resolve imports, APIs, and project conventions. Existing methods inject this knowledge as long i
arXiv:2606.05647v1 Announce Type: cross Abstract: AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to cod
arXiv:2606.05793v1 Announce Type: new Abstract: While LLM-based agents excel at individual tasks, effective collaboration with realistic human partners remains challenging. Most of the existing conver
arXiv:2603.00573v2 Announce Type: replace Abstract: Large language models (LLMs) achieve remarkable performance on diverse downstream and domain-specific tasks via parameter-efficient fine-tuning (PEF
arXiv:2606.05988v1 Announce Type: cross Abstract: Reasoning models produce long chain-of-thought traces that are costly to distill and encourage verbose student outputs. We study post-hoc compression
arXiv:2606.06022v1 Announce Type: new Abstract: Stance detection on social media is challenging due to short, noisy, and context-dependent language. While large language models (LLMs) show zero-shot g
arXiv:2606.05661v1 Announce Type: cross Abstract: Continual learning, the ability of AI systems to improve through sequential experience, has attracted substantial interest, but no high-quality benchm
Critical context on the new Anthropic blog: 1, AGI is *harder* than RSI (as used below). AGI: machine can do anything human can do, autonomously [not achieved] RSI (as used below): AI is a useful codi
arXiv:2606.06203v1 Announce Type: new Abstract: Input length and the position of relevant information are widely cited as the primary causes of degraded LLM long-context performance. Here, we study le
arXiv:2606.06217v1 Announce Type: new Abstract: When a disaster unfolds, responders must answer not only what is happening, but also why it is happening, what will happen next, and what to do now, oft
arXiv:2505.05026v5 Announce Type: replace Abstract: User interface (UI) design goes beyond visuals to shape user experience (UX), underscoring the shift toward UI/UX as a unified concept. While recent
arXiv:2508.15851v2 Announce Type: replace Abstract: Despite rapid progress in large language models (LLMs), current QA benchmarks still overlook the core challenge of real-world scientific information
arXiv:2606.05569v1 Announce Type: new Abstract: Mispronunciation Detection and Diagnosis (MDD) has gained increasing importance in computer-assisted language learning and speech technology in recent y
arXiv:2606.05233v1 Announce Type: cross Abstract: Recent computer-using-agent (CUA) red-teaming papers report prompt-injection attack success rates (ASR) of 42-98%, but these headline numbers cluster
arXiv:2601.21288v2 Announce Type: replace-cross Abstract: Autonomous driving is an important and safety-critical task, and recent advances in LLMs/VLMs have opened new possibilities for reasoning and
arXiv:2606.05179v1 Announce Type: new Abstract: Punctuation restoration improves ASR (Automatic Speech Recognition) readability. However streaming ASR requires online decisions with limited future con
arXiv:2605.24500v2 Announce Type: replace Abstract: This technical report presents our solution, EgoAdapt (Egocentric Adaptation via Category, Calibration, and Consistency), to the CVPR 2026 HD-EPIC V
arXiv:2606.06038v1 Announce Type: new Abstract: We study English-to-Prakrit machine translation in a low-resource setting where the target language is unsupported by IndicTrans2. We adapt the multilin
arXiv:2606.05168v1 Announce Type: new Abstract: Training on synthetic data causes model collapse, but existing analyses treat this as single-chain degradation. In reality, the AI ecosystem involves cr
arXiv:2606.05874v1 Announce Type: new Abstract: Current evaluations for Multimodal Large Language Models (MLLMs) overwhelmingly focus on utility-driven objectives, leaving model behavior under logic-n
even though @activegraphai is not a memory tool, it's a runtime built around memory so it can do well there cool to see a third party verify this and see that activegraph does well on privacy programs
arXiv:2511.10254v2 Announce Type: replace Abstract: Facial Emotion Analysis (FEA) extends traditional facial emotion recognition by incorporating explainable, fine-grained reasoning. The task integrat
arXiv:2606.05949v1 Announce Type: new Abstract: Scientific illustrations are essential tools for communicating research findings, especially in natural science, where they visualize complex concepts a
arXiv:2408.11336v3 Announce Type: replace-cross Abstract: Climate change stands as one of the most pressing global challenges of the twenty-first century, with far-reaching consequences such as rising
arXiv:2606.06211v1 Announce Type: new Abstract: Automatic speech recognition (ASR) has advanced remarkably for standard speech; however, pathological speech from neurological conditions remains a sign
arXiv:2606.06266v1 Announce Type: new Abstract: Hate speech detection is inherently subjective: people from different demographic groups perceive the same content very differently. Collecting enough a