CompanyAnthropic8 recent entries30 Apr 2026Risk Reporting for Developers' Internal AI Model UsearXiv:2604.24966v1 Announce Type: cross Abstract: Frontier AI companies first deploy their most advanced models internally, for weeks or months of safety testing, evaluation, and iteration, before a p→12 May 2026ConFit v3: Improving Resume-Job Matching with LLM-based Re-RankingarXiv:2605.09760v1 Announce Type: new Abstract: A reliable resume-job matching system helps a company find suitable candidates from a pool of resumes and helps a job seeker find relevant jobs from a l
CompanyOpenAI5 recent entries10 Apr 2026LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot InterfacesarXiv:2604.06188v1 Announce Type: cross Abstract: People increasingly hold sustained, open-ended conversations with large language models (LLMs). Public reports and early studies suggest that, in such→23 Apr 2026ESGLens: An LLM-Based RAG Framework for Interactive ESG Report Analysis and Score PredictionarXiv:2604.19779v1 Announce Type: new Abstract: Environmental, Social, and Governance (ESG) reports are central to investment decision-making, yet their length, heterogeneous content, and lack of stan→24 Jun 2026IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPOarXiv:2606.23032v2 Announce Type: replace Abstract: Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier language model→30 Jun 2026Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogatesarXiv:2606.30085v1 Announce Type: new Abstract: Large-language models have proven to be remarkable if inconsistent parrots of public attitudes and opinions. The extent to which LLMs are able to produc→7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr
CompanyGoogle7 recent entries29 May 2026FinVerBench: Benchmark Validity and Calibration in Large Language Model Financial Statement VerificationarXiv:2605.29586v1 Announce Type: new Abstract: We introduce FinVerBench, a benchmark and validity study for financial statement verification: determining whether a set of corporate financial statemen→24 Jun 2026IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPOarXiv:2606.23032v2 Announce Type: replace Abstract: Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier language model→30 Jun 2026PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM AgentsarXiv:2606.29225v1 Announce Type: new Abstract: LLM agents handle user requests on behalf of organizations through tool calls and must follow the company policies stated in their system prompts. Prior→9 Jul 2026From Content to Audience: A Multimodal Annotation Framework for Broadcast Television AnalyticsarXiv:2603.26772v2 Announce Type: replace-cross Abstract: Automated semantic annotation of broadcast television content presents distinctive challenges, combining structured audiovisual composition, d→28 Jul 2026Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models, and a Roster Instrument Captures 0.5% of ItarXiv:2607.23893v1 Announce Type: cross Abstract: Prior work on AI brand visibility measures the firm: does a model recommend a company, and does that track its reputation. This study asks the questio→3 Aug 2026Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model ReviewarXiv:2607.28631v1 Announce Type: new Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and compari→7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr
CompanyMeta8 recent entries14 Apr 2026Reasoning as Data: Representation-Computation Unity and Its Implementation in a Domain-Algebraic Inference EnginearXiv:2604.10908v1 Announce Type: new Abstract: Every existing knowledge system separates storage from computation. We show this separation is unnecessary and eliminate it. In a standard triple is_a(A→15 Apr 2026Fully Homomorphic Encryption on Llama 3 model for privacy preserving LLM inferencearXiv:2604.12168v1 Announce Type: cross Abstract: The applications of Generative Artificial Intelligence (GenAI) and their intersections with data-driven fields, such as healthcare, finance, transport→6 May 2026The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligencearXiv:2408.12622v3 Announce Type: replace-cross Abstract: Artificial intelligence (AI) is reshaping society, from video generation to medical diagnosis, coding agents to autonomous vehicles. Yet resea→12 May 2026V4FinBench: Benchmarking Tabular Foundation Models, LLMs, and Standard Methods on Corporate Bankruptcy PredictionarXiv:2605.10896v1 Announce Type: new Abstract: Corporate bankruptcy prediction is a high-stakes financial task characterized by severe class imbalance and multi-horizon forecasting demands. Public da→25 Jun 2026Small edits, large models: How Wikipedia advocacy shapes LLM valuesarXiv:2606.24890v1 Announce Type: new Abstract: Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia appears in near→9 Jul 2026From Content to Audience: A Multimodal Annotation Framework for Broadcast Television AnalyticsarXiv:2603.26772v2 Announce Type: replace-cross Abstract: Automated semantic annotation of broadcast television content presents distinctive challenges, combining structured audiovisual composition, d→3 Aug 2026Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model ReviewarXiv:2607.28631v1 Announce Type: new Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and compari→7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr
CompanyMistral1 recent entries7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr
CompanyxAI4 recent entries10 Apr 2026Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of InterestarXiv:2604.08525v1 Announce Type: cross Abstract: Today's large language models (LLMs) are trained to align with user preferences through methods such as reinforcement learning. Yet models are beginni→14 Apr 2026Dynamic Forecasting and Temporal Feature Evolution of Stock Repurchases in Listed Companies Using Attention-Based Deep Temporal NetworksarXiv:2604.09650v1 Announce Type: cross Abstract: Accurately predicting stock repurchases is crucial for quantitative investment and risk management, yet traditional static models fail to capture the →28 Jul 2026Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models, and a Roster Instrument Captures 0.5% of ItarXiv:2607.23893v1 Announce Type: cross Abstract: Prior work on AI brand visibility measures the firm: does a model recommend a company, and does that track its reputation. This study asks the questio→7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr
CompanyDeepSeek4 recent entries1 Jun 2026Learning Whom to Trust: Market-Feedback Adaptive Retrieval for Frozen LLMs in Event-Driven Financial RAGarXiv:2605.31201v1 Announce Type: new Abstract: Financial retrieval-augmented generation (RAG) systems typically rank evidence by textual relevance, but in financial markets the useful evidence source→6 Jun 2026Can LLMs Write Correct TLA+ Specifications? Evaluating Natural-Language-to-TLA+ GenerationarXiv:2606.05792v1 Announce Type: new Abstract: TLA+ has supported industrial verification at companies such as Amazon and Microsoft, yet writing correct TLA+ specifications from natural language stil→10 Jun 2026Instruction Finetuning DeepSeek-R1-8B Model Using LoRA and NEFTunearXiv:2606.10392v1 Announce Type: new Abstract: Financial named-entity recognition (NER) is essential for translating unstructured financial reports and news into structured knowledge graphs. However,→30 Jun 2026Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogatesarXiv:2606.30085v1 Announce Type: new Abstract: Large-language models have proven to be remarkable if inconsistent parrots of public attitudes and opinions. The extent to which LLMs are able to produc