JaiTTS: A Thai Voice Cloning Model
arXiv:2604.27607v1 Announce Type: new Abstract: We present JaiTTS-v1.0, a state-of-the-art Thai voice cloning text-to-speech model built through continual training on a large Thai-centric speech corpu
Knowledge catalogue
arXiv:2604.27607v1 Announce Type: new Abstract: We present JaiTTS-v1.0, a state-of-the-art Thai voice cloning text-to-speech model built through continual training on a large Thai-centric speech corpu
arXiv:2604.27661v1 Announce Type: new Abstract: Detecting language ideologies is a valuable yet complex task for understanding how identities are constructed through discourse. In Luxembourg's multicu
arXiv:2604.27998v1 Announce Type: cross Abstract: Latent reasoning offers a more efficient alternative to explicit reasoning by compressing intermediate reasoning into continuous representations and s
arXiv:2604.27039v1 Announce Type: new Abstract: Token serves as the fundamental unit of computation in modern autoregressive models, and generation length directly influences both inference cost and r
arXiv:2604.27712v1 Announce Type: cross Abstract: Scene-text image captioning requires fusing three information streams -- visual features, OCR-detected text, and linguistic knowledge -- to generate d
arXiv:2502.14541v3 Announce Type: replace Abstract: The rapid advancement of Large Language Models (LLMs) has opened new opportunities in recommender systems by enabling zero-shot recommendation witho
arXiv:2604.27345v1 Announce Type: new Abstract: Human annotators frequently disagree on emotion labels, yet most evaluations of Large Language Model (LLM) emotion annotation collapse these judgments i
arXiv:2509.15549v2 Announce Type: replace Abstract: Multilingual instruction fine-tuning (IFT) empowers large language models to generalize across diverse linguistic and cultural contexts; however, hi
arXiv:2604.28061v1 Announce Type: cross Abstract: Numerous metascience studies and other initiatives have begun to monitor the prevalence of open science practices when it is more important to underst
arXiv:2604.27393v1 Announce Type: new Abstract: Recent progress in multimodal large language models (MLLMs) has brought AI capabilities from static offline data processing to real-time streaming inter
arXiv:2604.28031v1 Announce Type: new Abstract: When researchers iteratively refine ideas with large language models, do the models preserve fidelity to the original objective? We introduce DriftBench
arXiv:2604.27846v1 Announce Type: new Abstract: How people narrate their experiences offers a window into how the mind organizes them. Computational approaches to therapeutic writing have evolved from
arXiv:2504.02768v4 Announce Type: replace Abstract: We introduce MultiBLiMP 1.0, a massively multilingual benchmark of linguistic minimal pairs, covering 101 languages and 2 types of subject-verb agre
arXiv:2604.28147v1 Announce Type: new Abstract: Surprisal theory links human processing effort to the predictability of an upcoming linguistic unit, but empirical work often leaves the notion of a uni
arXiv:2604.27401v1 Announce Type: new Abstract: Perturbation probing generates task-specific causal hypotheses for FFN neurons in large language models using two forward passes per prompt and no backp
arXiv:2601.11908v2 Announce Type: replace Abstract: Large language models (LLMs) struggle with reasoning over long contexts where relevant information is sparsely distributed. Although plan-and-execut
arXiv:2604.27379v1 Announce Type: new Abstract: Dialogue models are inherently reactive, responding to the current user turn without anticipating upcoming intents, which leads to redundant interaction
arXiv:2604.27533v1 Announce Type: new Abstract: Evaluating automatic speech recognition (ASR) systems is a classical but difficult and still open problem, which often boils down to focusing only on th
arXiv:2505.19630v4 Announce Type: replace Abstract: Large language models (LLMs) struggle in real-world clinical consultations. Single-turn consultation systems require patients to describe all sympto
arXiv:2604.27850v1 Announce Type: new Abstract: Task-based dialogue systems assist users in achieving specific goals, such as executing actions or retrieving information, through natural language inte
arXiv:2604.27616v1 Announce Type: new Abstract: People commonly leverage structured content to accelerate knowledge acquisition and research problem solving. Among these, roadmaps guide researchers th
arXiv:2604.27467v1 Announce Type: cross Abstract: Code sandboxes have emerged as a critical infrastructure for advancing the coding capabilities of large language models, providing verifiable feedback
arXiv:2604.27204v1 Announce Type: new Abstract: In the field of universal automatic phonetic transcription (APT), clean and diverse training transcriptions are required. However, such high-quality dat
arXiv:2604.27169v1 Announce Type: new Abstract: We show that the geometric relations between semantic features in large language models' hidden states closely mirror human psychological associations.
arXiv:2604.27439v1 Announce Type: new Abstract: This study analyzes Indonesian student opinions on the adoption of artificial intelligence in higher education using two approaches: TF-IDF-based machin
arXiv:2604.27488v1 Announce Type: new Abstract: We introduce Skills-Coach, a novel automated framework designed to significantly enhance the self-evolution of skills within Large Language Model (LLM)-
arXiv:2604.28048v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used as proxies for human perception in urban analysis, yet it remains unclear whether persona prompting p
arXiv:2602.16516v2 Announce Type: replace Abstract: This paper introduces ParlaCAP, a large-scale dataset for analyzing parliamentary agenda setting across Europe, and proposes a cost-effective method
arXiv:2604.27468v1 Announce Type: new Abstract: Maintaining information in context is essential in successful real-time language comprehension, but maintenance is cognitively costly and can slow proce
arXiv:2604.27232v1 Announce Type: new Abstract: Models of sign language have historically lagged behind those for spoken language (text and speech). Recent work has greatly improved their performance
arXiv:2604.27296v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for code editing, yet the prevalent full-code generation paradigm suffers from severe efficiency bo
arXiv:2604.27861v1 Announce Type: cross Abstract: Decompositional jailbreaks pose a critical threat to large language models (LLMs) by allowing adversaries to fragment a malicious objective into a seq
arXiv:2604.28021v1 Announce Type: cross Abstract: Cooking is a cultural expression of human creativity that transcends geography and time through the orchestration of ingredients and techniques, much
arXiv:2506.05412v3 Announce Type: replace-cross Abstract: Where someone looks is a nonverbal communication cue that children and adults readily use. How well can Vision-Language Models (VLMs) infer ga
arXiv:2508.13024v3 Announce Type: replace Abstract: LLM-based web agents have the potential to automate long-running web tasks, such as searching for products in multiple e-shops and subsequently orde
arXiv:2604.27398v1 Announce Type: new Abstract: For constructing text embeddings, mean pooling, which averages token embeddings, is the standard approach. This paper examines whether mean pooling actu
arXiv:2602.02980v2 Announce Type: replace-cross Abstract: In this work, we focus on front-end design for speech deepfake detectors, the component that determines the discriminative acoustic cues provi
arXiv:2604.27844v1 Announce Type: cross Abstract: Communication has emerged as a critical bottleneck in the distributed training of large language models (LLMs). While numerous approaches have been pr
arXiv:2604.26351v1 Announce Type: new Abstract: Language models (LMs) behave more like humans when their cognitive resources are restricted, particularly in predicting sentence processing costs such a
arXiv:2604.26230v1 Announce Type: new Abstract: I developed a new version of Latent Semantic Scaling (LSS) employing word2vec as a masked language model. Unlike original spatial models, it assigns pol
arXiv:2604.25979v1 Announce Type: cross Abstract: We present a quantitative analysis of character-pair substitution ratios in the Voynich manuscript, testing whether Currier's A/B language distinction
arXiv:2604.26319v1 Announce Type: new Abstract: Stance detection identifies the attitude of a text author toward a given target. Recent studies have explored various LLM-based strategies for this task
arXiv:2604.26779v1 Announce Type: cross Abstract: RL post-training of frontier language models is increasingly bottlenecked by autoregressive rollout generation, making rollout acceleration a central
arXiv:2603.16496v2 Announce Type: replace Abstract: Large language model (LLM) agents increasingly rely on external memory to support long-horizon interaction, personalized assistance, and multi-step
arXiv:2604.26326v1 Announce Type: cross Abstract: Reinforcement learning (RL) has unlocked complex reasoning abilities in large language models (LLMs). However, most RL algorithms suffer from performa
arXiv:2604.00706v2 Announce Type: replace Abstract: Assessing the veracity of a claim made online is a complex and important task with real-world implications. When these claims are directed at commun
arXiv:2601.17617v3 Announce Type: replace-cross Abstract: LLM-powered search agents are increasingly being used for multi-step information seeking tasks, yet the IR community lacks empirical understan
arXiv:2604.25931v1 Announce Type: new Abstract: We identify a previously unknown calibration property of large language models: providing one confirmed intermediate fact toward a multi-step reasoning
arXiv:2604.25930v1 Announce Type: new Abstract: We study whether a structured recurrent state can serve as a compact associative backbone for language modeling while still supporting exact retrieval.
arXiv:2604.26310v1 Announce Type: new Abstract: Fine-grained emotion classification, which identifies specific emotional states such as happiness, anger, sadness, and fear, remains a challenging task
arXiv:2604.26148v1 Announce Type: cross Abstract: AI agents operating on user interfaces must understand how interfaces communicate state and feedback to act reliably. As a core communicative modality
arXiv:2604.26048v1 Announce Type: new Abstract: This paper presents a principled and scalable framework for systematically generating complex Question Answering (QA) data. In the core of this framewor
arXiv:2604.26176v1 Announce Type: cross Abstract: The integration of Large Language Models (LLMs) with Retrieval-Augmented Generation (RAG) has significantly advanced Knowledge Graph Question Answerin
arXiv:2604.26923v1 Announce Type: cross Abstract: LLMs have achieved strong results on both function-level code synthesis and repository-level code modification, yet a capability that falls between th
arXiv:2604.26312v1 Announce Type: new Abstract: Public opinion towards the Free Nutritious Meal Program (MBG) on YouTube social media reflects diverse community responses. This study applies the Long
arXiv:2604.25928v1 Announce Type: new Abstract: Professional domain knowledge underpins human civilization, serving as both the basis for industry entry and the core of complex decision-making and pro
arXiv:2604.26229v1 Announce Type: new Abstract: This study compares machine learning and deep learning approaches for cyberbullying detection in Indonesian-language Instagram comments. Using a balance
arXiv:2506.13116v2 Announce Type: replace-cross Abstract: Crime hotspot prediction is critical for ensuring urban safety and effective law enforcement, it remains challenging due to complex spatial de
arXiv:2506.20876v4 Announce Type: replace Abstract: Technological progress has led to concrete advancements in tasks that were regarded as challenging, such as automatic fact-checking. Interest in ado
arXiv:2604.26768v1 Announce Type: new Abstract: Parametric Retrieval-Augmented Generation (PRAG) encodes external documents into lightweight parameter modules that can be retrieved and merged at infer