NSFA = "Not-Secure-For-Agents" — Ant Group's guardrail family for agent operational security, a different axis from the content-safety SingGuard line: a CIA-triad taxonomy of 185 agentic risk variants (prompt injection, data exfiltration, unsafe tool actions), served in dual mode — a generative audit model and a ~50ms classifier — with claimed >94% F1. Four text models (0.8B/2B/4B/9B on Qwen3.5, Apache-2.0) plus the NSFA_Benchmarks suite (~97K samples across 133 languages). Technical report not yet published ("arXiv link coming soon").

Outputs 2

SingGuard-NSFA models

model
License Apache 2.0
Base model qwen3.5

Variants

Name Parameters Notes
SingGuard-NSFA-0.8B 0.8B
SingGuard-NSFA-2B 2B
SingGuard-NSFA-4B 4B
SingGuard-NSFA-9B 9B

NSFA_Benchmarks

dataset
Size ~97K samples, 133 languages
safetyguardrailsagentsopen-weight

Related