AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

Content type
87,171Total entries
1Added by human
87,170Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
51,191 results
Safety

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

DGX agent

arXiv:2606.11709v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with the distri

safetyarxiv-cs-cl
11 Jun 2026
Research
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SDQM: Synthetic Data Quality Metric for Object Detection Dataset Evaluation

DGX agent

arXiv:2510.06596v2 Announce Type: replace-cross Abstract: The performance of machine learning models depends heavily on training data. The scarcity of large-scale, well-annotated datasets poses signif

researcharxiv-cs-ai
11 Jun 2026
Model Releases

Sparsified Kolmogorov-Arnold Networks for Interpretable Quantum State Tomography

DGX agent

arXiv:2606.11814v1 Announce Type: cross Abstract: Machine-learning approaches to quantum state tomography can achieve high reconstruction fidelity, but the physical structure used by the trained model

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

System Report for CCL25-Eval Task 5: New Dataset and LoRA-Fine-Tuned Qwen2.5

DGX agent

arXiv:2606.12392v1 Announce Type: cross Abstract: Recently, large language models (LLMs) have achieved promising progress in the fields of classical Chinese translation and the generation of classical

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

The N-Body Problem: Parallel Execution from Single-Person Egocentric Video

DGX agent

arXiv:2512.11393v2 Announce Type: replace Abstract: Humans can intuitively parallelise complex activities, but can a model predict this from observing a single person? Given one egocentric video, we i

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

Toward Trustworthy AI: Multi-Target Adversarial Attacks and Robust Defenses for Continuous Data Summarization

DGX agent

arXiv:2606.11804v1 Announce Type: new Abstract: Trustworthy AI requires reliable data-processing pipelines, not only robust downstream predictive models. As an upstream component, data summarization d

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping

DGX agent

arXiv:2502.11034v3 Announce Type: replace Abstract: Loss spikes remain a persistent obstacle in large-scale language model pretraining. While previous research has attempted to identify the root cause

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness

DGX agent

arXiv:2606.10577v1 Announce Type: new Abstract: Zero-shot vision-and-language navigation in continuous environments (VLN-CE) has recently become feasible with large vision-language models (VLMs). Howe

model-releasesarxiv-cs-ro
10 Jun 2026
Research

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval

DGX agent

arXiv:2606.10657v1 Announce Type: new Abstract: Multiple-choice (MCQA) benchmarks are the standard for evaluating pretrained large language models, but their reliance on log-likelihood scoring makes t

researcharxiv-cs-cl
10 Jun 2026
Model Releases

Benchmarking Knowledge Editing using Logical Rules

DGX agent

arXiv:2606.10554v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications that require access to up-to-date knowledge. However, retraining LLM

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

ChartLens: A Dual-Branch Framework for Chart Data Correction and Factual Summary Refinement

DGX agent

arXiv:2606.10640v1 Announce Type: new Abstract: In this report, we present our champion solution for the DataMFM Challenge Track 2: Chart Understanding. This track requires models to recover structure

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

CodeAlchemy: Synthetic Code Rewriting at Scale

DGX agent

arXiv:2606.10087v1 Announce Type: new Abstract: Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats. While synthetic data has proven transformative f

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

Evaluating Research-Level Math Proofs via Strict Step-Level Verification

DGX agent

arXiv:2606.10799v1 Announce Type: new Abstract: Large Language Models (LLMs) struggle to rigorously verify complex mathematical proofs. Standard global evaluation approaches suffer from 'context poiso

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

LLM-Based Code Documentation Generation and Multi-Judge Evaluation

DGX agent

arXiv:2606.09852v1 Announce Type: cross Abstract: High-quality source code documentation is vital yet often neglected, especially in critical domains like healthcare where reliability and maintainabil

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Mind the Gap: Can Frontier LLMs Pass a Standardized Office Proficiency Exam?

DGX agent

arXiv:2606.10956v1 Announce Type: new Abstract: The deployment of Large Language Model (LLM) agents for computer automation is accelerating, yet their ability to navigate complex, professional-grade p

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

MMClima: A Framework for Multimodal Climate Science Data and Evaluation

DGX agent

arXiv:2606.10194v1 Announce Type: cross Abstract: Climate change research increasingly requires AI systems that reason across text, dynamic visual content, and scientific figures, yet existing climate

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization

DGX agent

arXiv:2606.10768v1 Announce Type: cross Abstract: The success of Large Language Models in mathematical reasoning relies heavily on the generation of diverse and valid solution paths during the rollout

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

READER: Robust Evidence-based Authorship Decoding via Extracted Representations

DGX agent

arXiv:2606.10794v1 Announce Type: new Abstract: As agentic applications increasingly route user tasks through official and third-party LLM APIs, provenance becomes an operational question: which model

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning

DGX agent

arXiv:2606.10254v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved near-perfect performance in solving high-school mathematics, their ability to evaluate the diverse reas

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

SHAPE: Coalition-Aware Expert Pruning for Sparse Mixture-of-Experts LLMs

DGX agent

arXiv:2606.09886v1 Announce Type: cross Abstract: Sparse Mixture-of-Experts (MoE) large language models achieve strong quality with low per-token compute, yet their deployment is often limited by the

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Small Data, Big Noise: Adversarial Training for Robust Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2606.10610v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) has become essential for adapting foundation models to downstream NLP tasks. However, current PEFT methods often

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios

DGX agent

arXiv:2606.10394v1 Announce Type: new Abstract: Large language models are increasingly used to power personal agents for everyday applications, but evaluating these agents remains a challenge. Existin

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

DGX agent

arXiv:2606.11070v1 Announce Type: cross Abstract: Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However,

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

The Order Matters: Sequential Fine-Tuning of LLaMA for Coherent Automated Essay Scoring

DGX agent

arXiv:2606.10327v1 Announce Type: new Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.g., lead, claim, evidence, conclusion), yet most approaches treat

model-releasesarxiv-cs-cl
10 Jun 2026
Research

Towards Robust Arabic Speech Emotion Recognition with Deep Learning

DGX agent

arXiv:2606.10278v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) aims to identify a speaker's emotional state from audio signals. While recent advances in deep learning have signific

researcharxiv-cs-ai
10 Jun 2026
Model Releases

TRAPS: Therapeutic Response Analysis via Pathway-informed Stratification

DGX agent

arXiv:2606.09898v1 Announce Type: new Abstract: Cancer treatment planning requires decisions across multiple clinical dimensions at once. Clinicians must determine whether a patient should receive tar

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

A Comparison of SSL-Based Feature Extractors and Back-End Classifiers for Spoofing Detection: A Multi-Corpus Training and Cross-Linguistic Analysis

DGX agent

arXiv:2606.08669v1 Announce Type: cross Abstract: Voice biometric systems face growing threats from spoofing attacks, yet the evaluation of detection models remains inconsistent across datasets. To in

model-releasesarxiv-cs-lg
9 Jun 2026
Safety

Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks

DGX agent

arXiv:2604.01039v2 Announce Type: replace-cross Abstract: System Instructions in Large Language Models (LLMs) are commonly used to enforce safety policies, define agent behavior, and protect sensitive

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems

DGX agent

arXiv:2606.07805v1 Announce Type: new Abstract: The rapid evolution of Large Language Models (LLMs) from passive assistants to autonomous, execution-capable agents has introduced critical operational

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Can Global XAI Methods Reveal Injected Behaviours in LLMs? SHAP vs Rule Extraction vs RuleSHAP

DGX agent

arXiv:2505.11189v3 Announce Type: replace Abstract: Large language models (LLMs) can amplify misinformation, undermining societal goals such as the UN SDGs. We study three documented drivers of misinf

model-releasesarxiv-cs-ai
9 Jun 2026
Research

Capacity, Not Format: Rethinking Structured Reasoning Failures

DGX agent

arXiv:2606.09410v1 Announce Type: new Abstract: Prior work treats structured output as a reasoning tax, but this framing is incomplete: the cost of formatting depends strongly on a model's spare capac

researcharxiv-cs-ai
9 Jun 2026
Model Releases

CATPO: Critique-Augmented Tree Policy Optimization

DGX agent

arXiv:2606.08346v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a dominant paradigm for improving the reasoning capabilities of large language models

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures

DGX agent

arXiv:2606.08275v1 Announce Type: cross Abstract: When an LLM agent fails -- issues a refund it should not have, calls the wrong tool, leaks data -- existing tooling answers what happened (observabili

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

CLASP: Language-Driven Robot Skill Selection and Composition using Task-Parameterized Learning

DGX agent

arXiv:2606.08169v1 Announce Type: cross Abstract: Enabling robots to understand and execute tasks from natural language commands while maintaining data efficiency remains challenging. Foundation model

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

CTS-Bench: Benchmarking Graph Coarsening Trade-offs for GNNs in Clock Tree Synthesis

DGX agent

arXiv:2602.19330v2 Announce Type: replace Abstract: Graph Neural Networks (GNNs) are increasingly explored for physical design analysis in Electronic Design Automation, particularly for modeling Clock

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

Curvature-Guided LoRA: Matching Full Fine-Tuning in Function Space

DGX agent

arXiv:2603.29824v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning methods such as LoRA enable efficient adaptation of large pretrained models, but often lag behind full fine-tuning i

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

Data Synthesis and Parameter-Efficient Fine-Tuning for Low-Resource NMT: A Case Study on Q'eqchi' Mayan

DGX agent

arXiv:2606.09767v1 Announce Type: cross Abstract: Neural machine translation for digitally low-resource Indigenous languages is often hindered by extreme data scarcity, prompting reliance on extractiv

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

De novo molecular generation with optical property preconditioning at the token level

DGX agent

arXiv:2606.08221v1 Announce Type: new Abstract: Designing OLED molecules with targeted optical properties remains challenging due to the scarcity of high-quality data and the limited reliability of co

model-releasesarxiv-cs-lg
9 Jun 2026
Safety

Diffuse AI Control on Fuzzy Tasks

DGX agent

arXiv:2606.08892v1 Announce Type: new Abstract: AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment. Diffuse AI Control is a subfiel

safetyarxiv-cs-lg
9 Jun 2026
Safety

Enhancing AI Interpretability and Safety through Localised Architectures

DGX agent

arXiv:2606.07998v1 Announce Type: cross Abstract: Recent advances in generative AI, especially powerful Large Language Models (LLMs) and Large Reasoning Models (LRMs), raise concerns over the interpre

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Evaluating Advanced Prompting on Gemini Flash for Multi-Hop Biomedical QA

DGX agent

arXiv:2606.07548v1 Announce Type: cross Abstract: The MedHopQA challenge presents a critical test for Large Language Models (LLMs): complex, multi-hop reasoning in the high-stakes biomedical domain. T

model-releasesarxiv-cs-ai
9 Jun 2026
Tutorials

Explaining Data Mixing Scaling Laws

DGX agent

arXiv:2606.08167v1 Announce Type: cross Abstract: Recent research has established empirical scaling laws to predict model performance on multi-domain data mixtures. However, a theoretical understandin

tutorialsarxiv-cs-ai
9 Jun 2026
Safety

Few-step Cofolding with All-Atom Flow Maps

DGX agent

arXiv:2606.08375v1 Announce Type: new Abstract: All-atom generative modeling of 3D biomolecular complexes has emerged as the dominant paradigm for predicting the structure of proteins and protein-liga

safetyarxiv-cs-lg
9 Jun 2026
Research

Forecasting Japanese elections: A nonlinear machine-learning approach

DGX agent

arXiv:2606.07572v1 Announce Type: cross Abstract: Despite Japan being one of the world's largest advanced democracies, the development of election forecasting models for its national elections remains

researcharxiv-cs-lg
9 Jun 2026
Model Releases

From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing

DGX agent

arXiv:2606.08932v1 Announce Type: cross Abstract: Rule-following agents tasked with executing policies and regulations often fail via Silent Scope Omission (SSO): a model applies a general rule but si

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

DGX agent

arXiv:2606.08530v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background s

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Generalized Rank-based Evaluation for Knowledge Graph Completion: Perspectives, Framework, and Analyses

DGX agent

arXiv:2606.08921v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to predict missing facts from an observed knowledge graph (KG), playing a crucial role in a wide range of real-wor

safetyarxiv-cs-lg
9 Jun 2026
Model Releases

Harnessing Streaming Video in the Wild

DGX agent

arXiv:2606.08615v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are increasingly required to process unbounded video streams in applications such as video-call assistants, live commentar

model-releasesarxiv-cs-cv
9 Jun 2026
← Previous
1…354355356357358…1067
Next →