AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries90,223
  • Agents7,699
  • Applications5,506
  • Concepts5
  • Hardware1,889
  • Industry6,186
  • Local Ai5,045
  • Model Releases24,499
  • Research20,615
  • Safety13,633
  • Syntheses17
  • Tools1,677
  • Tutorials3,452

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries90,223
  • Agents7,699
  • Applications5,506
  • Concepts5
  • Hardware1,889
  • Industry6,186
  • Local Ai5,045
  • Model Releases24,499
  • Research20,615
  • Safety13,633
  • Syntheses17
  • Tools1,677
  • Tutorials3,452

Source
HumanDGX agent

Content type
AllBlog
90,223Total entries
1Added by human
90,222Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,101 results
Safety

SemStruct: Contextualizing Semantic Embeddings with Structural Information for Schema Matching

DGX agent

arXiv:2605.30729v1 Announce Type: new Abstract: Schema matching is a fundamental step in integrating heterogeneous data sources. While Pre-trained Language Models (PLMs) have revolutionized this task

safetyarxiv-cs-lg
1 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Tutorials

Spectral Reach: Understanding Neural Scaling as Progress into the Spectral Tail

DGX agent

arXiv:2605.31244v1 Announce Type: new Abstract: Neural scaling laws describe predictable power-law relationships between model size, dataset size, compute, and performance. While these laws guide the

tutorialsarxiv-cs-lg
1 Jun 2026
Model Releases

The Surface You Test Is Not the Surface That Breaks

DGX agent

arXiv:2605.30454v1 Announce Type: cross Abstract: Tool-augmented LLM agents are vulnerable to prompt injection: a third party who controls part of the agent's context can plant instructions that the a

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

UniDial-EvalKit: A Unified Toolkit for Evaluating Multi-Faceted Conversational Abilities

DGX agent

arXiv:2603.23160v2 Announce Type: replace Abstract: Benchmarking large language models (LLMs) and agents in multi-turn interactive scenarios is essential for understanding their practical capabilities

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification

DGX agent

arXiv:2509.24901v4 Announce Type: replace-cross Abstract: Although probing frozen models has become a standard evaluation paradigm, self-supervised learning in audio defaults to fine-tuning when pursu

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Active Learning for Machine Learning Driven Molecular Dynamics

DGX agent

arXiv:2509.17208v3 Announce Type: replace Abstract: Machine-learned coarse-grained (CG) potentials are fast, but degrade over time when simulations reach under-sampled bio-molecular conformations, and

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

AgentCVR: Active Multi-Agent Cross-Video Reasoning via Script-Simulated Reinforcement Learning

DGX agent

arXiv:2605.29643v1 Announce Type: new Abstract: Cross-Video Reasoning (CVR) has emerged as a critical frontier in multimodal intelligence, requiring models to retrieve, align, and aggregate evidence d

model-releasesarxiv-cs-cv
29 May 2026
Industry

Agree! was talking about this with @havoyan just a few days ago. That's also the reason why so much of the value has been accruing to the fr…

DGX agent

Agree! was talking about this with @havoyan just a few days ago. That's also the reason why so much of the value has been accruing to the frontier models in my opinion (cc @GavinSBaker) because if you

industryclem-delangue--x
29 May 2026
Model Releases

Beyond English and Evasion: A Human-Annotated Multi-Domain Benchmark for High-Stakes LLM Safety Evaluation in Chinese

DGX agent

arXiv:2605.29667v1 Announce Type: new Abstract: When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

CalArena: A Large-Scale Post-Hoc Calibration Benchmark

DGX agent

arXiv:2605.30188v1 Announce Type: cross Abstract: Reliable probability estimates are critical in many machine learning applications, yet modern classifiers are often poorly calibrated. Post-hoc calibr

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Connecting Independently Trained Modes via Layer-Wise Connectivity

DGX agent

arXiv:2505.02604v5 Announce Type: replace Abstract: Empirical studies have shown that continuous low-loss paths can be constructed between independently trained neural network models. This phenomenon,

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Deep Adaptive Dimension Reduction for Bayesian Inference in Inverse Problems

DGX agent

arXiv:2605.29373v1 Announce Type: new Abstract: Solving high-dimensional PDE-governed inverse problems is often challenging due to complex non-Gaussian posterior distributions, expensive forward model

model-releasesarxiv-cs-lg
29 May 2026
Tutorials

Deep Psychovisual Image Representations

DGX agent

arXiv:2605.29260v1 Announce Type: new Abstract: Psychovisual models suggest human vision decouples low-level feature extraction from higher cognition by first forming intermediate abstractions. In con

tutorialsarxiv-cs-cv
29 May 2026
Model Releases

DynSess: Dynamic Session-Level Evaluation and Optimization Framework for Role-Playing Agents

DGX agent

arXiv:2605.29256v1 Announce Type: cross Abstract: Role-playing with large language models is fundamentally a session-level task, requiring agents to sustain character identity and interaction quality

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

ESPO: Early-Stopping Proximal Policy Optimization

DGX agent

arXiv:2605.29860v1 Announce Type: cross Abstract: When a large language model under reinforcement learning commits a wrong reasoning step early in a trajectory, standard algorithms force it to keep ge

model-releasesarxiv-cs-ai
29 May 2026
Research

From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges

DGX agent

arXiv:2601.08654v2 Announce Type: replace-cross Abstract: Rubric-based text evaluation increasingly uses large language models (LLMs) as scalable judges, but aligning frozen black-box models with huma

researcharxiv-cs-ai
29 May 2026
Safety

GAP3D: Generative Alignment of VLM Latents to Patch-Level Embeddings for 3D Generation

DGX agent

arXiv:2605.28995v1 Announce Type: new Abstract: Recent approaches integrating vision-language models (VLMs) as prompt encoders for generative model conditioning typically rely on expensive end-to-end

safetyarxiv-cs-cv
29 May 2026
Model Releases

GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization

DGX agent

arXiv:2605.29107v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rank products, documents, and recommendations for user queries, which makes manipulating these rankings a gr

model-releasesarxiv-cs-ai
29 May 2026
Tutorials

Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence

DGX agent

arXiv:2605.30093v1 Announce Type: new Abstract: Foundation features from self-supervised vision models and text-to-image diffusion models have proven effective for semantic correspondence estimation.

tutorialsarxiv-cs-cv
29 May 2026
Model Releases

GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents

DGX agent

arXiv:2605.29668v1 Announce Type: new Abstract: LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

How Braintrust turns customer requests into code with Codex

DGX agent

Braintrust leverages OpenAI's Codex model to automatically convert customer requests and natural language specifications into functional code, streamlining the software development process. This appli

model-releasesopenai
29 May 2026
Model Releases

How Far Ahead Do LLMs Plan? Uncovering the Latent Horizon in Chain-of-Thought Reasoning

DGX agent

arXiv:2602.02103v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) reasoning has become a central mechanism for eliciting multi-step reasoning in Large Language Models (LLMs). Yet recent

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

LiveSVG: Zero-Shot SVG Animation via Video Generation

DGX agent

arXiv:2605.30174v1 Announce Type: new Abstract: We introduce LiveSVG, a zero-shot approach for generating Scalable Vector Graphics (SVG) animations using video diffusion models. Current SVG animation

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs

DGX agent

arXiv:2605.29737v1 Announce Type: cross Abstract: LLM-based coding assistants are seeing rapid adoption, offering substantial gains in developer productivity. As organizations increasingly ship code t

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions

DGX agent

arXiv:2605.29738v1 Announce Type: cross Abstract: Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs

DGX agent

arXiv:2605.29300v1 Announce Type: cross Abstract: Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses ar

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMs

DGX agent

arXiv:2605.29716v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising non-autoregressive generative paradigm. Given the prohibitive computational cost of

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs

DGX agent

arXiv:2605.29685v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly applied in social contexts such as emotional companionship and customer service, measuring their social

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems

DGX agent

arXiv:2605.29676v1 Announce Type: new Abstract: Large language models in Agentic AI systems consume tool schemas and execution results and emit tool invocations as structured data. The default languag

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields

DGX agent

arXiv:2605.29833v1 Announce Type: new Abstract: As multimodal language models play an increasingly important role in scientific research, materials science offers a critical testbed due to its interdi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

On the Construction and Implications of Low-Loss Valleys in LoRA-based Bayesian Inference

DGX agent

arXiv:2605.29580v1 Announce Type: new Abstract: While parameter-efficient fine-tuning methods like low-rank adaptation (LoRA) are standard for large language models, principled estimation of epistemic

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Overcoming Forgetting in LLM Fine-Tuning with Evolution Strategies

DGX agent

arXiv:2605.30148v1 Announce Type: cross Abstract: Evolution Strategies (ES) has recently emerged as a competitive alternative to reinforcement learning (RL) for large language model (LLM) fine-tuning,

model-releasesarxiv-cs-ai
29 May 2026
Safety

PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

DGX agent

arXiv:2605.29582v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown promise as educational tutors, yet effective tutoring requires more than solving problems: it must provide pro

safetyarxiv-cs-cl
29 May 2026
Model Releases

Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning

DGX agent

arXiv:2605.29028v1 Announce Type: cross Abstract: Conditioned Sequence Models (CSMs) learn policies by treating return-to-go (RTG) as a control signal. However, existing CSMs often treat the RTGs as s

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Robust and Efficient Guardrails with Latent Reasoning

DGX agent

arXiv:2605.29068v1 Announce Type: new Abstract: Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrai

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

DGX agent

arXiv:2605.29468v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to support scientific work, but it is unclear whether they uphold responsible conduct of research (

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge

DGX agent

arXiv:2605.29402v1 Announce Type: cross Abstract: Understanding long-form egocentric videos remains challenging for multimodal large language models (MLLMs) due to limited context length and insuffici

model-releasesarxiv-cs-ai
29 May 2026
Safety

Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning

DGX agent

arXiv:2605.29032v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss. However, powerful RL optimizers inevitably

safetyarxiv-cs-lg
29 May 2026
Model Releases

Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection

DGX agent

arXiv:2605.30344v1 Announce Type: new Abstract: Recent advances in Vision-Language Models (VLMs) have achieved impressive performance across many tasks, yet prior studies report unsatisfactory perform

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Took me a while to figure out what all the ESMFold2 rage was about. At first, the benchmarking data didn't look super remarkable to me but i…

DGX agent

Took me a while to figure out what all the ESMFold2 rage was about. At first, the benchmarking data didn't look super remarkable to me but it turns there are many impressive aspects: - Fully open sour

model-releasesyann-lecun--x
29 May 2026
Model Releases

TRACE: Toulmin-based Reasoning Assessment through Constructive Elements for LLM CoT Evaluation

DGX agent

arXiv:2605.29656v1 Announce Type: new Abstract: Evaluating open-ended outputs from large language models (LLMs) remains challenging due to the absence of ground truth. Existing metrics rely on final-a

model-releasesarxiv-cs-ai
29 May 2026
Safety

When and How Long? The Readout-Mediator Angle in Temporal Reasoning

DGX agent

arXiv:2605.29126v1 Announce Type: cross Abstract: A linear probe can decode a representation almost perfectly and yet be completely irrelevant to how the model uses it. On calendar-date duration reaso

safetyarxiv-cs-ai
29 May 2026
Model Releases

Adaptive Reservoir Computing for Multi-Scenario Chaotic System Forecasting

DGX agent

arXiv:2605.28145v1 Announce Type: new Abstract: We present an adaptive reservoir computing framework for the CTF-4-Science Lorenz benchmark, which evaluates machine learning models across twelve disti

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications

DGX agent

arXiv:2605.27761v1 Announce Type: new Abstract: The rapid development of GUI foundation models and mobile GUI agents has spurred numerous evaluation benchmarks, yet most rely on simulated environments

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Beyond being fast, LiteParse is designed to provide highly accurate, semantically coherent text for LLM use. We benchmarked every open-sourc…

DGX agent

Beyond being fast, LiteParse is designed to provide highly accurate, semantically coherent text for LLM use. We benchmarked every open-source, model-free PDF parser on LLM QA tasks - from PyPDF to PyM

model-releasesjerry-liu--x
28 May 2026
Model Releases

Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

DGX agent

arXiv:2510.05291v2 Announce Type: replace Abstract: As Large Language Models (LLMs) develop stronger multilingual capabilities, their sensitivity to culturally diverse entities becomes increasingly im

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

ClinicalAgents: Multi-Agent Orchestration for Clinical Decision Making with Dual-Memory

DGX agent

arXiv:2603.26182v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have demonstrated potential in healthcare, they often struggle with the complex, non-linear reasoning required fo

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Cost-Sensitive Evaluation for Binary Classifiers

DGX agent

arXiv:2510.22016v2 Announce Type: replace Abstract: Selecting an appropriate evaluation metric for classifiers is crucial for model comparison, parameter optimization, and deployment decisions, yet th

model-releasesarxiv-cs-lg
28 May 2026
← Previous
1…454455456457458…1357
Next →