Virasai AI / VaHive Systems Lab

The package you approved isn’t the one running tomorrow.

MCP servers update. Dependencies shift, tool schemas widen, install scripts appear, descriptions your model reads as instructions get rewritten. Sentinel records exactly what an npm artifact declares — without ever running it — and tells you what changed between two releases.

Apache-2.0 and complete. No withheld tier.

Field note 01 / The system is the unit of analysis Scroll to examine the layers
The problem, concretely

Trust is earned at version 1.2 and spent at version 1.7.

Almost every update is routine. The one that isn’t looks identical from the outside — same package, same name, same install command — and registries make versions discoverable without telling you what changed inside one.

An agent makes this worse than an ordinary dependency. It reads tool descriptions as instructions, and it acts on them without a human in the loop. Something has to be recording the shape of each artifact as it arrives, or the release that matters passes unnoticed.

approved baselinepackage@1.2.0 / artifact recorded
evidence-linked changedependency + tool schema difference
deliberate next stepreview / accept / freeze / investigate
Why this is trustworthy

We publish where our own tools fail.

Sentinel’s limitations page reports two numbers, not one: across a pinned corpus of 50 real published MCP servers, 37 yield a usable tool inventory and only 12 can be resolved completely. Both are stated because only the second permits the conclusion that a tool was removed. It names the five specific ways extraction breaks, and warns about the most likely way to misread a report. The corpus is checked in with exact versions and digests, so you can re-run the figure rather than take it on trust.

complete, always

Change detection

Artifact digest, per-file inventory, dependencies, install scripts and entrypoints are recorded on every package, every time. If something moved, you will know.

inferred, stated

Tool surface

Recovered by parsing shipped JavaScript. Of 50 corpus packages, 37 yield a usable inventory and 12 resolve completely; the rest say why not. A list without complete: true is a lower bound.

never

Execution

Nothing is run, imported or started. No install scripts, no MCP session. A package can do things at runtime that no report mentions, and we say so.

not our call

The verdict

Findings are facts with a severity input. Sentinel does not decide whether a package is malicious, and an empty findings list does not mean clean.

Also open, also free

The tools we build are public before they are products.

Public repository

Magus OpenSecMCP

A local Rust MCP security gateway that sits between an MCP client and downstream tool servers.

Inspect the repository
Early / public repository

Sentinel

Deterministic, non-executing evidence and change monitoring for public npm MCP servers.

View Sentinel on GitHub
The wider picture

Different controls solve different parts of the problem.

A policy written in prose does not enforce itself. A model can make nuanced assessments, but it cannot make a deterministic claim. A guardrail can help, but it should not be asked to be the whole system.

A practical control map

A single “AI safety layer” is not a serious model of the system.

  1. 01

    Human objective & operating authority

    The decision context: what a person or organisation intends to authorise.

  2. 02

    Agent model

    A probabilistic system that generates, reasons, classifies, and proposes actions.

  3. 03

    Guardrails & evaluators

    Useful screening and routing controls; valuable, but not a universal guarantee.

  4. 04

    Policy & permissions

    The explicit rules defining who or what may act, and under which conditions.

  5. 05

    Deterministic execution boundary

    Repeatable specified checks or blocks before tools, APIs, files, or other actions.

  6. 06

    Evidence & review over time

    Audit records, version history, and human judgement as context, authority, and dependencies change.

Time is a control surface too: permissions, dependencies, context, and accumulated risk can change after an agent is deployed.

Published preprints

We state the research record—and its limits—plainly.

These papers frame questions and proposed architectures. They are not product certification or evidence that a theoretical mechanism has been deployed.

Zenodo / 14 March 2026

MAGUS v3.0: A Governance Architecture for Structural Alignment Drift in Long-Running Agentic AI Systems

A theoretical governance architecture and open issues register for structural alignment drift in long-running agentic systems. It is not a report of a deployed product.

Read the preprint ↗
Zenodo / 19 May 2026

Connecting Activation Geometry to Execution Intent: A Multi-Representation Framework for Detecting Computational Divergence in Agentic AI

A proposed framework for comparing input, expressed reasoning, activation-level representations, and execution intent. It makes no empirical performance claims.

Read the preprint ↗
Work with us

Need a rigorous read of an agent system before it becomes a problem?

We offer structural governance and drift audits, adversarial issues registers, and scoped advisory work. The same boundaries apply: we say what we can assess before an engagement begins.

Choose a route