Test-Time Curriculum for Open-Set AIGC Detection
arXiv:2608.00559v1 Announce Type: new Abstract: AI-generated image detectors deployed in open-world environments inevitably face distribution shifts as new and stronger generative models continue to e
Knowledge catalogue
arXiv:2608.00559v1 Announce Type: new Abstract: AI-generated image detectors deployed in open-world environments inevitably face distribution shifts as new and stronger generative models continue to e
arXiv:2608.00916v1 Announce Type: cross Abstract: Modern reranking recipes---billion-scale cross-encoders, mixture-of-experts (MoE) backbones, and distillation against strong teachers---have outpaced
arXiv:2608.02050v1 Announce Type: new Abstract: Can a strictly local, iterated, weight-shared computation primitive support language modelling, and which of those three properties actually drives the
arXiv:2608.02588v1 Announce Type: cross Abstract: In [AS21], Axiotis and Sviridenko conjectured that the linear dependence on the restricted condition number in sparse convex optimization cannot be im
// The confidence cliff in self-improving autoresearch // Autoresearch loops are still quite brittle. Here is a nice paper offering some insights into why this might be happening. Self-improving autor
arXiv:2608.00507v1 Announce Type: new Abstract: Perceptual narrowing---the developmental loss of non-native phoneme discrimination in the first year of life itep{werker1984}---is a canonical developme
arXiv:2608.02306v1 Announce Type: new Abstract: We introduce a mathematical framework for shape comparison based on mapping functions from the shape domain to a common reference domain. This Push-Forw
arXiv:2608.02138v1 Announce Type: new Abstract: Current speech translation systems, including SpeechLLMs, are trained on cleaned text and tend to strip disfluencies like filled pauses and false starts
Sam Sabin / Axios: The UK AISI says it observed a total of 19 instances where Mythos and GPT-5.6 Sol tried to hack people and companies during a routine cyber evaluation in July — The U.K. AI Security
this post from 3 years ago and commercial LLMs *still* can’t play chess anywhere near as well serious players (except by calling external tools) When @GaryMarcus and others point out that GPT-4 is bad
Together AI gives developers a high-throughput production path for running DeepSeek V4 Flash across coding, tool-use, and agentic workloads. Start building: https://www.together.ai/models/deepseek-v4-
arXiv:2603.08091v2 Announce Type: replace Abstract: Large language model (LLM)-based judges are widely adopted for automated evaluation and reward modeling, yet their judgments are often affected by j
arXiv:2606.18621v2 Announce Type: replace Abstract: Relational databases are widely used for managing structured data in real-world systems. Detecting anomalies from such relational data is crucial fo
arXiv:2505.09710v4 Announce Type: replace Abstract: We investigate deep morphological neural networks (DMNNs), studying how changes in algebraic structure affect the expressivity and trainability of d
arXiv:2608.01284v1 Announce Type: new Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the
arXiv:2608.00640v1 Announce Type: new Abstract: Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high
arXiv:2608.00573v1 Announce Type: cross Abstract: Serving Mixture-of-Experts (MoE) large language models across distributed edge servers is bottlenecked by the cross-server expert transmission. The ex
arXiv:2608.00042v1 Announce Type: new Abstract: Domain adaptation of small language models (SLMs) has emerged as a practical strategy for deploying capable NLP systems in resource-constrained, high-st
arXiv:2608.01833v1 Announce Type: cross Abstract: Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generali
arXiv:2608.01471v1 Announce Type: new Abstract: Understanding sentiment in low-resource languages remains a key challenge for Natural Language Processing (NLP), particularly when domain-specific data
arXiv:2608.01518v1 Announce Type: new Abstract: Continual learning in clinical imaging faces a dual challenge: a model must assimilate knowledge from new anatomical domains while retaining representat
arXiv:2608.02528v1 Announce Type: new Abstract: Mixtures of low-rank adaptation experts increase parameter-efficient capacity by routing each input through a subset of adapters. Recent dynamic routers
arXiv:2608.00915v1 Announce Type: new Abstract: Uplift modeling (conditional-average-treatment-effect estimation) drives personalized targeting, yet published uplift benchmarks frequently disagree on
What's Changed Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model's MTP head for speculative decoding automatically /v1/chat/completions streaming now matches OpenAI's wire format: rol
arXiv:2608.00094v1 Announce Type: new Abstract: Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a
arXiv:2608.02598v1 Announce Type: new Abstract: Aerial-ground person re-identification is a challenging task due to cross-platform viewpoint variations, which cause severe occlusion and geometric defo
We're detailing two new incidents that occurred during external cyber evaluations conducted by independent evaluation partners. We outline what happened, how the activity was contained, and how we’re
arXiv:2608.00105v1 Announce Type: new Abstract: Pathology foundation models are reported to encode molecular programmes in tissue morphology, but the evidence is usually a cohort-wide ranked gene list
arXiv:2603.06591v2 Announce Type: replace-cross Abstract: Transformers frequently allocate disproportionate attention to specific tokens, a phenomenon known as attention sinks. Causal large language m
arXiv:2608.00013v1 Announce Type: new Abstract: Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fu
arXiv:2608.01602v1 Announce Type: new Abstract: Cardiac digital twins convert clinical images into physiological measurements through observation operators, yet calibration studies often assume a fixe
arXiv:2608.01409v1 Announce Type: new Abstract: Biomedical fact-checking systems must do more than predict whether a claim is supported, contradicted, or unaddressed: they should also produce evidence
arXiv:2608.01004v1 Announce Type: new Abstract: Platform teams hosting agent-extensibility surfaces face a regression-economics paradox: every onboarding customer ships an evaluation set tuned to thei
I use A LOT both openAI and Anthropic products. When I need some frontend work (pure web dev) (or answer that feel less verbose and more to the point) I use Anthropic. For multimodality openAI feels b
It's more accurate to say that many highly engaged online gamers are hostile to AI, not that 'gamers' as a whole are. Gaming is a huge community with hundreds of millions of people, and opinions vary
arXiv:2608.01388v1 Announce Type: cross Abstract: Runtime safety monitors based on Linear Temporal Logic (LTL) and finite automata (FSA) are increasingly deployed to intercept unsafe tool-call sequenc
arXiv:2608.02412v1 Announce Type: new Abstract: Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the
arXiv:2608.02603v1 Announce Type: new Abstract: Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the appa
arXiv:2608.00582v1 Announce Type: new Abstract: Pretrained byte-level BPE tokenizers can segment underrepresented languages inefficiently. Replacing a tokenizer changes the meaning of nearly every tok
arXiv:2608.00036v1 Announce Type: new Abstract: Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that
arXiv:2603.28568v2 Announce Type: replace Abstract: Vision-language models (VLMs) share visual-textual representations across zero-shot classification, image captioning, and visual question answering
arXiv:2608.02271v1 Announce Type: new Abstract: Parameter-Efficient Fine-tuned (PEFT) models are frequently downloaded from open repositories by practitioners. This widespread practice creates a signi
arXiv:2605.06340v2 Announce Type: replace-cross Abstract: Continuous post-deployment compliance audits, mandated by emerging regulations such as the EU AI Act and Digital Services Act, create a class
arXiv:2607.29122v1 Announce Type: new Abstract: Pixel-space diffusion models aim to learn an end-to-end generator directly over raw pixels. This is challenging because a single model must capture both
arXiv:2607.28968v1 Announce Type: cross Abstract: Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a st
arXiv:2607.28858v1 Announce Type: cross Abstract: Automatic brain tumor segmentation from magnetic resonance imaging (MRI) has become a fundamental task in computer-assisted diagnosis, treatment plann
arXiv:2604.02330v2 Announce Type: replace-cross Abstract: Recent advances in video diffusion have enabled the development of 'world models' capable of simulating interactive environments. However, the
arXiv:2509.22310v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has achieved impressive results across domains, yet learning an optimal policy typically requires extensive intera
AFAIK the most significant breakthrough since 2017 besides scaling old ideas was broadening from base models into larger systems that incorporate symbol/manipulating entities like harnesses, tools, an
arXiv:2607.29626v1 Announce Type: new Abstract: As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly i
arXiv:2603.20075v2 Announce Type: replace-cross Abstract: Compilers are critical to modern computing, yet fixing compiler bugs is difficult. While recent large language model (LLM) advancements enable
AI9Stars has released G9v3-39A5B an open weights language model designed to deliver even stronger reasoning capabilities than ai9stars/G9v3-3B with its 39B and 5 active experts. It is released under t
Alibaba Group Holding Ltd. today debuted a new addition to its Qwen series of open-source large language models. Qwen3.8-Max is the Chinese e-commerce giant’s most capable LLM to date. It features 2.4
All frontier AI models (including GPT-5.6 Sol, Fable 5, and Kimi K3) consistently fail to implement solvers for basic nonlinear PDEs correctly, often introducing significant errors in numerical stabil
arXiv:2607.29549v1 Announce Type: new Abstract: Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers remains challen
An internal version of our next major model produced 10 new results on long-standing open problems in mathematics and theoretical computer science, using roughly $2,000 worth of tokens at GPT-5.6 Sol
arXiv:2607.29456v1 Announce Type: cross Abstract: Double Machine Learning (DML) is a popular approach for treatment effect estimation in various settings, which allows a wide range of flexible machine
Qwen3.8-Max from Alibaba’s Qwen team achieved second place in the Vision Arena benchmark, scoring 1,305 points. It trails only Claude Fable 5 (High), which leads by a slim 13‑point margin. The post un
arXiv:2607.29539v1 Announce Type: cross Abstract: Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has
arXiv:2607.28661v1 Announce Type: new Abstract: Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding