Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
22 commits
Select commit Hold shift + click to select a range
153b8a3
fix(ingest): inspect unverified OMS bundle contents
yashrajp22 Oct 7, 2026
7d95757
test(ingest): retain normal findings for unsigned OMS fixtures
yashrajp22 Oct 7, 2026
45f5c04
Merge branch 'main' into yashraj/scan-unverified-oms-content
github-actions[bot] Oct 7, 2026
2712948
Merge branch 'main' into yashraj/scan-unverified-oms-content
github-actions[bot] Oct 8, 2026
88de383
Merge branch 'main' into yashraj/scan-unverified-oms-content
github-actions[bot] Oct 8, 2026
14f6a36
Merge branch 'main' into yashraj/scan-unverified-oms-content
github-actions[bot] Oct 8, 2026
f0d1d14
Merge branch 'main' into yashraj/scan-unverified-oms-content
github-actions[bot] Oct 8, 2026
ac214bf
fix(oms): inspect decoded bundle content without binary false positives
yashrajp22 Oct 9, 2026
52d03b3
fix(oms): preserve readable Unicode in binary projections
yashrajp22 Oct 9, 2026
6674959
test(oms): assert detection independently of risk thresholds
yashrajp22 Oct 9, 2026
24d87bd
test(oms): exercise unknown base64 in SC3 eligible carrier
yashrajp22 Oct 9, 2026
6ecc53d
fix(oms): keep escaped duplicate-key bundles explicitly partial
yashrajp22 Oct 9, 2026
c43480d
style(test): format OMS regression parameters
yashrajp22 Oct 9, 2026
52a3f78
Merge branch 'main' into yashraj/scan-unverified-oms-content
github-actions[bot] Oct 9, 2026
871c00e
Merge branch 'main' into yashraj/scan-unverified-oms-content
github-actions[bot] Oct 9, 2026
dbf8453
Merge branch 'main' into yashraj/scan-unverified-oms-content
github-actions[bot] Oct 9, 2026
f1fbee6
Merge branch 'main' into yashraj/scan-unverified-oms-content
github-actions[bot] Oct 9, 2026
09a357f
Merge branch 'main' into yashraj/scan-unverified-oms-content
github-actions[bot] Oct 9, 2026
e60a2b2
merge: preserve OMS projections with Python source classification
yashrajp22 Oct 9, 2026
bd96815
merge: preserve security fixes with latest review branch
yashrajp22 Oct 9, 2026
820ac92
Merge branch 'main' into yashraj/scan-unverified-oms-content
github-actions[bot] Oct 9, 2026
a2385ee
Merge branch 'main' into yashraj/scan-unverified-oms-content
github-actions[bot] Oct 9, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 10 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -915,13 +915,15 @@ SkillSpector uses a two-stage detection pipeline:
- High recall (catches most issues)
- Moderate precision (some false positives)

A valid, root-level OpenSSF Model Signing signature (`skill.oms.sig`) is retained in the
component inventory as type `oms_signature`, but excluded from static and LLM content analysis.
OMS bundles necessarily contain long base64-encoded payload, signature, and certificate fields;
generic obfuscated-code checks can otherwise misclassify those fields as hidden executable content.
The recognizer checks the minimal OMS DSSE/in-toto structure; it does not verify the signature,
certificate chain, transparency-log entry, or signer identity. Invalid or unrecognized signature
files are scanned normally.
OpenSSF Model Signing / DSSE bundles remain in content analysis, regardless of
filename. Static and LLM analyzers inspect decoded JSON payloads, readable DER
certificate/signature content, and all unknown wrapper fields. Strictly parsed
binary fields are projected rather than treated as executable base64; canonical
bytes remain available to YARA and fingerprints. Unsupported encodings (including
environments without the optional `cryptography` dependency) retain raw fields
and produce an explicit incomplete-analysis warning. Findings in decoded content
point to the original bundle and label their projection in evidence. None of this
verifies signatures, certificate chains, transparency logs, or signer identity.

### Stage 2: LLM Semantic Analysis (Optional)
- Evaluates context and intent
Expand Down Expand Up @@ -950,7 +952,7 @@ The tool requires outbound HTTPS access to `api.osv.dev` for live vulnerability
SkillSpector is defense-in-depth, not a sandbox. Know what it does and does not do before relying on it:

- **It never executes the scanned skill.** All analysis is static (regex, Python AST, YARA) plus optional LLM evaluation of file *contents* — the skill's code is never run.
- **LLM analysis sends analyzer-eligible file contents to the configured provider.** When LLM analysis is enabled (the default), file contents are sent to the active `SKILLSPECTOR_PROVIDER` endpoint. Recognized OMS signature files are excluded. Use `--no-llm` to keep contents local (static analysis only).
- **LLM analysis sends analyzer-eligible file contents to the configured provider.** When LLM analysis is enabled (the default), file contents are sent to the active `SKILLSPECTOR_PROVIDER` endpoint. DSSE payloads and readable signing metadata are included. Use `--no-llm` to keep contents local (static analysis only).
- **SC4 sends dependency names to OSV.dev.** The supply-chain check queries [OSV.dev](https://osv.dev) with the package names and versions the skill declares, to look up known CVEs. This is fundamental to the check and runs even with `--no-llm`. It sends dependency coordinates (not file contents), requires no API key, and falls back to a bundled list when OSV.dev is unreachable.
- **It does not sandbox the host.** SkillSpector flags risky patterns *before* you install a skill; it does not contain or isolate a skill you choose to install anyway.

Expand Down
2 changes: 1 addition & 1 deletion docs/DEVELOPMENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,7 +81,7 @@ All targets assume the virtual environment is **already created and activated**.
| `zip_bytes`, `mode` | Optional zip input and scan mode |
| `components` | List of relative file paths in the skill |
| `file_cache` | Map of path → file contents |
| `inspection_ledger` | Structured evidence for files excluded, skipped, or failed during analysis; a recognized OMS signature is recorded as an `oms_signature` scope exclusion. |
| `inspection_ledger` | Structured evidence for files excluded, skipped, or failed during analysis; unsupported DSSE decoding is recorded as an `oms_signature` partial event; successfully projected bundles remain in analysis without a scope exclusion. |
| `ast_cache` | Map of path → AST representation (for future use) |
| `manifest`, `previous_manifest` | Parsed skill metadata (e.g. from SKILL.md) |
| `component_metadata` | List of dicts: path, type, lines, executable, size_bytes (from build_context) |
Expand Down
2 changes: 1 addition & 1 deletion src/skillspector/inspection_ledger.py
Original file line number Diff line number Diff line change
Expand Up @@ -154,7 +154,7 @@ class LedgerReason(StrEnum):
LedgerReason.MANIFEST_ABSENT: ("No compatible manifest was present for this analyzer."),
LedgerReason.NO_APPLICABLE_FILES: ("No files matched this analyzer's applicability contract."),
LedgerReason.OMS_SIGNATURE: (
"Recognized OMS signature metadata is excluded from content analysis."
"DSSE content decoding is incomplete or unsupported; raw fields were retained for inspection. Signature authenticity was not verified."
),
LedgerReason.BASELINE_FILE: (
"The explicitly selected suppression baseline is excluded from content analysis."
Expand Down
38 changes: 38 additions & 0 deletions src/skillspector/nodes/analyzers/static_runner.py
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,7 @@
observe_analyzer_findings,
)
from skillspector.nodes.deduplicate import classification_metadata_key
from skillspector.oms import project_oms_content
from skillspector.python_ast import (
MAX_PYTHON_AST_SOURCE_CHARS,
ParsedPythonFile,
Expand Down Expand Up @@ -2310,9 +2311,46 @@ def _scan_all_views_detailed(
started_at: float | None = None,
python_ast: ParsedPythonFile | None = None,
python_source: bool | None = None,
_oms_projected: bool = False,
) -> tuple[list[Finding], LedgerReason | None, dict[str, int | float]]:
"""Scan bounded raw windows and return any limit with observed/limit metrics."""
started_at = time.monotonic() if started_at is None else started_at
runtime_limit = MAX_STATIC_ANALYSIS_SECONDS_PER_ARTIFACT
if timeout_seconds is not None:
runtime_limit = min(runtime_limit, max(0.0, timeout_seconds))
now = time.monotonic()
if now >= started_at + runtime_limit:
return (
[],
LedgerReason.RUNTIME_LIMIT,
{
"observed_seconds": max(0.0, now - started_at),
"limit_seconds": runtime_limit,
},
)
projection = None if _oms_projected else project_oms_content(content)
if projection is not None and projection.view.name != "raw":
findings, reason, metrics = _scan_all_views_detailed(
path,
projection.view.text,
pattern_modules,
None,
max_findings=max_findings,
timeout_seconds=timeout_seconds,
started_at=started_at,
# This view is derived JSON metadata, not a Python execution surface.
python_source=False,
_oms_projected=True,
)
for finding in findings:
finding.start_line = finding.end_line = 1
finding.start_column = finding.end_column = None
finding.evidence["oms_projection"] = (
"Decoded content in the original DSSE bundle; line 1 identifies the carrier."
)
finding.evidence[_SOURCE_START_EVIDENCE] = projection.view.source_offset(0)
finding.evidence.pop(_SOURCE_END_EVIDENCE, None)
return findings, reason, metrics
ast_modules = [module for module in pattern_modules if _uses_python_ast(module)]
lexical_modules = [module for module in pattern_modules if not _uses_python_ast(module)]
if python_source is None:
Expand Down
59 changes: 39 additions & 20 deletions src/skillspector/nodes/build_context.py
Original file line number Diff line number Diff line change
Expand Up @@ -73,6 +73,7 @@
is_executable_content,
is_zip_content,
)
from skillspector.oms import project_oms_content
from skillspector.python_ast import (
PythonSourceClassification,
classify_python_source,
Expand Down Expand Up @@ -833,7 +834,7 @@ def _decode_base64_json(value: object) -> dict[str, object] | None:


def _is_valid_oms_signature_bytes(data: bytes) -> bool:
"""Recognize the minimal OMS DSSE/in-toto structure from bounded bytes."""
"""Recognize OMS structure only; this does not verify or trust its contents."""
Comment thread
yashrajp22 marked this conversation as resolved.
try:
if len(data) > MAX_FILE_BYTES:
return False
Expand Down Expand Up @@ -2806,16 +2807,7 @@ def _record_processing_runtime(*, phase: str, path: str, now: float) -> None:
elif signature_valid:
recognized_oms_signature_paths.add(_OMS_SIGNATURE_PATH)
recognized_oms_signatures = frozenset(recognized_oms_signature_paths)
signature_events = [
ledger_event(
outcome=LedgerOutcome.OUT_OF_SCOPE,
record_type=LedgerRecordType.SCOPE_BOUNDARY,
phase="discovery",
path=path,
reason=LedgerReason.OMS_SIGNATURE,
)
for path in sorted(recognized_oms_signatures)
]
signature_events: list[InspectionLedgerEvent] = []
baseline_events = [
ledger_event(
outcome=LedgerOutcome.OUT_OF_SCOPE,
Expand All @@ -2826,12 +2818,6 @@ def _record_processing_runtime(*, phase: str, path: str, now: float) -> None:
)
for path in sorted(selected_baselines)
]
for artifact in artifact_inventory:
if artifact["path"] in recognized_oms_signatures:
artifact["disposition"] = ArtifactDisposition.OUT_OF_SCOPE
artifact["reason"] = LedgerReason.OMS_SIGNATURE.value
llm_file_cache.pop(artifact["path"], None)

primary_path = next(
(path for path in ("SKILL.md", "skill.md") if path in inventoried_components), None
)
Expand Down Expand Up @@ -2999,8 +2985,6 @@ def _record_processing_runtime(*, phase: str, path: str, now: float) -> None:
ordinary_components: list[str] = []
for path in cache_candidates:
ordinary_artifact = inventory_by_path.get(path)
if path in recognized_oms_signatures:
continue
if path in raw_file_cache or (
ordinary_artifact is not None
and ordinary_artifact.get("disposition")
Expand Down Expand Up @@ -3321,7 +3305,7 @@ def mark_excluded_nested_metadata(
]
)
)
for path in [*recognized_containers, *recognized_oms_signatures]:
for path in recognized_containers:
Comment thread
yashrajp22 marked this conversation as resolved.
llm_file_cache.pop(path, None)

postprocessing_events: list[InspectionLedgerEvent] = []
Expand Down Expand Up @@ -3494,6 +3478,41 @@ def _mark_runtime_partial(affected_paths: list[str], first_limited_path: str) ->
)
)

for path, content in local_file_cache.items():
if not content.lstrip().startswith("{"):
continue
projection_started = monotonic()
if projection_started >= processing_deadline:
_record_processing_runtime(
phase="signature_projection", path=path, now=projection_started
)
continue
projection = project_oms_content(content)
if projection is None:
continue
projection_finished = monotonic()
if projection_finished >= processing_deadline:
_record_processing_runtime(
phase="signature_projection", path=path, now=projection_finished
)
continue
if path in llm_file_cache:
llm_file_cache[path] = projection.view.text
if not projection.complete:
signature_events.append(
ledger_event(
outcome=LedgerOutcome.PARTIAL,
record_type=LedgerRecordType.SYSTEM,
phase="discovery",
path=path,
reason=LedgerReason.OMS_SIGNATURE,
)
)
artifact = inventory_by_path.get(path)
if artifact is not None and artifact["disposition"] == ArtifactDisposition.ANALYZED:
artifact["disposition"] = ArtifactDisposition.PARTIAL
artifact["reason"] = LedgerReason.OMS_SIGNATURE.value

llm_components = sorted(llm_file_cache)
file_cache = dict(llm_file_cache)
manifest_events: list[InspectionLedgerEvent] = []
Expand Down
171 changes: 171 additions & 0 deletions src/skillspector/oms.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,171 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0

"""Bounded, untrusted DSSE content projections; never signature verification."""

from __future__ import annotations

import base64
import binascii
import json
import re
from array import array
from dataclasses import dataclass

from skillspector.artifacts import SecurityTextView

_MAX_CHARS = 1_000_000
_MAX_NODES = 10_000
_PRINTABLE = re.compile(r"[^\x00-\x1f\x7f-\x9f\ufffd]{4,}")


@dataclass(frozen=True)
class OMSProjection:
view: SecurityTextView
complete: bool


def _object(pairs: list[tuple[str, object]]) -> dict[str, object]:
result: dict[str, object] = {}
for key, value in pairs:
if key in result:
raise ValueError("duplicate JSON key")
result[key] = value
return result


def _load(text: str) -> object:
value = json.loads(text, object_pairs_hook=_object)
pending = [(value, 0)]
count = 0
while pending:
item, depth = pending.pop()
count += 1
if count > _MAX_NODES or depth > 64:
raise ValueError("JSON projection limit")
if isinstance(item, dict):
pending.extend((child, depth + 1) for child in item.values())
elif isinstance(item, list):
pending.extend((child, depth + 1) for child in item)
return value


def _decode(value: object) -> bytes:
if not isinstance(value, str):
raise ValueError("expected base64 string")
return base64.b64decode(value, validate=True)


def _printable(data: bytes) -> list[str]:
# Preserve readable Unicode and format characters for semantic/obfuscation
# analysis. Invalid bytes and control bytes separate runs rather than
# silently joining attacker-controlled fragments.
return _PRINTABLE.findall(data.decode("utf-8", errors="replace"))


def project_oms_content(content: str) -> OMSProjection | None:
"""Decode DSSE payloads independently of file names or media-type spelling.

Only strictly parsed DER certificate/signature fields lose their base64
representation. Their readable content and every unknown wrapper field stay
in the view. Unsupported fields stay raw and make coverage partial. Neither
successful decoding nor a complete projection establishes signer trust.
"""
if not content.lstrip().startswith("{") or len(content) > _MAX_CHARS:
return None
try:
bundle = _load(content)
except (ValueError, RecursionError):
recognized = "dsseEnvelope" in content
if not recognized:
# Recognition only: a duplicate-key bundle may escape its field
# names. Never use this permissive parse as analysis content.
try:
untrusted = json.loads(content)
recognized = isinstance(untrusted, dict) and "dsseEnvelope" in untrusted
except (ValueError, RecursionError):
pass
return OMSProjection(SecurityTextView("raw", content), False) if recognized else None
if not isinstance(bundle, dict) or "dsseEnvelope" not in bundle:
return None
envelope = bundle["dsseEnvelope"]
if not isinstance(envelope, dict):
return OMSProjection(SecurityTextView("raw", content), False)
complete = True
try:
payload = _load(_decode(envelope.get("payload")).decode("utf-8"))
if not isinstance(payload, (dict, list)):
raise ValueError("expected JSON payload")
envelope["payload"] = payload
except (ValueError, UnicodeError, binascii.Error, RecursionError):
complete = False

# cryptography is optional: environments without it inspect the raw binary
# fields and explicitly report the interpretation gap.
try:
from cryptography import x509
from cryptography.exceptions import UnsupportedAlgorithm
from cryptography.hazmat.primitives.asymmetric.utils import (
decode_dss_signature,
encode_dss_signature,
)
from cryptography.hazmat.primitives.serialization import Encoding

material = bundle.get("verificationMaterial")
chain = material.get("x509CertificateChain") if isinstance(material, dict) else None
certificates = chain.get("certificates") if isinstance(chain, dict) else None
if not isinstance(certificates, list) or not 0 < len(certificates) <= 64:
complete = False
else:
for entry in certificates:
try:
if not isinstance(entry, dict):
raise ValueError("expected certificate object")
der = _decode(entry.get("rawBytes"))
cert = x509.load_der_x509_certificate(der)
if cert.public_bytes(Encoding.DER) != der:
raise ValueError("noncanonical certificate")
readable = {
"subject": cert.subject.rfc4514_string(),
"issuer": cert.issuer.rfc4514_string(),
"extensions": [str(extension.value) for extension in cert.extensions],
"printable_der": _printable(der),
}
entry["rawBytes"] = readable
except (
ValueError,
TypeError,
binascii.Error,
x509.DuplicateExtension,
UnsupportedAlgorithm,
):
complete = False
signatures = envelope.get("signatures")
if not isinstance(signatures, list) or not 0 < len(signatures) <= 64:
complete = False
else:
for entry in signatures:
try:
if not isinstance(entry, dict):
raise ValueError("expected signature object")
der = _decode(entry.get("sig"))
r, s = decode_dss_signature(der)
if (
not (0 < r.bit_length() <= 521 and 0 < s.bit_length() <= 521)
or encode_dss_signature(r, s) != der
):
raise ValueError("noncanonical ECDSA signature")
entry["sig"] = {"printable_der": _printable(der)}
except (ValueError, TypeError, binascii.Error):
complete = False
except ImportError:
complete = False

projected = json.dumps(bundle, ensure_ascii=False, indent=2)
if len(projected) > _MAX_CHARS:
return OMSProjection(SecurityTextView("raw", content), False)
# Findings refer to the original carrier, not synthetic files or lines in
# decoded data. Their evidence explicitly labels this derived view.
return OMSProjection(
SecurityTextView("oms-decoded", projected, array("I", [0]) * len(projected)), complete
)
Loading