A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety Classification
arXiv cs.AIen
arXiv:2608.21570v1 Announce Type: new Abstract: Deploying a safety layer for large language models on commodity hardware is constrained by the guards available to do it: current open guard models hold between 1 and 9 billion parameters, are oriented toward the graphics processing unit, and answer in seconds per request on a central processing unit. This paper presents a reproducible, license-aware knowledge-distillation recipe addressing that constraint. A strong open guard labels a corpus of roughly 97,000 prompts, drawn from 24 public datasets, into seven safety categories aligned to a public hazard taxonomy, and a fleet of small students spanning lexical, shallow, encoder and generative a
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- RIACT: A Responsible AI System for Personalized Study Habit Tracking and Early Burnout Signal Detection in University StudentsarXiv cs.AI · August 25, 2026
- Retrieval-grounded robot program generation and simulation-based correction via Model Context ProtocolarXiv cs.AI · August 25, 2026
- Quantifying geographic domain shift to decouple the geospatial transferability of human mobility flow generation modelsarXiv cs.AI · August 25, 2026
- Data-Driven Dynamic Algorithm Dispatch with Large Language ModelsarXiv cs.AI · August 25, 2026
- Generate in the Chart, Not on the Boundary: Function-Symbol Grounding for Hard Constraints in LTN-GANsarXiv cs.AI · August 25, 2026
- From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student SimulationarXiv cs.AI · August 25, 2026