Research record

METHOD PUBLISHED WITH EVERY RESULT

ALGIERS, DZ · UTC+01:00

Everything below is measurement-only, authorized, or conducted on samples already in hand. Where a finding is unpublished it is marked IN PROGRESS and is not available for sale, download or redistribution.

ACTIVE

Capability sandboxing for on-device agents

Containment design for a system that lets a model execute code: bind-mount discipline, declared-permission plugins, a network allowlist enforced beneath the application rather than inside it, and escape testing aimed at the boundary itself. Includes a self-audit that confirmed a privileged bridge crossing the guest boundary, the guards that held against cross-tree ptrace and proc re-rooting, and the fixes that closed the gap.

Method
Sandboxing · Permission Models · Network Enforcement · Boundary Testing
Finding
Privileged bridge (Shizuku) crosses guest boundary; proc filters block cross-tree ptrace; native-offload is the escape vector
Reproduction
sandbox-escape-report.md
Target
Escape-test suite passing against Shizuku bridge + native-offload vector
ETA
2026-10-15
ACTIVE

Degradation behaviour in long-horizon agents

Observed failure modes when a model runs unattended across many turns: text-repetition collapse, empty responses returned as success, hard length walls, tool loops, and truncated context silently ending a task. Each classified as a distinct failure with an explicit exit rather than a retry that burns budget and pretends nothing happened.

Method
Failure Analysis · Loop Detection · Run Instrumentation
Finding
7 distinct failure classes; TextRepetitionGuard, EmptyResponseGuard, EOFStubGuard, LengthWallGuard, VerificationStopGuard, BudgetResumeGuard, SemaphoreConcurrencyGuard
Reproduction
Minis X agent loop — run telemetry in /var/minis/runs/
Target
7 guard classes instrumented + telemetry dashboard
ETA
2026-10-20
ACTIVE

Sustaining a platform fork

Operational study of carrying a divergent platform fork: on-device build verification instead of trusting CI, expectation-gated releases, a debt ledger of files too large to keep editing safely, and a triage protocol for deciding what to steal from other forks. The failure modes are more instructive than the wins.

Method
Release Discipline · Technical Debt · CI Contracts
Finding
CI green ≠ correct; on-device verification is the only ground truth; debt ledger prevents god-file accumulation
Reproduction
Minis X CI pipeline + on-device verification protocol
Target
Expectation-gated release automation + debt ledger CI check
ETA
2026-10-25
ACTIVE

Abstention as a first-class retrieval output

A retrieval system rebuilt so "the corpus does not contain this" is a common and correct answer. Keyed lookup data and conceptual corpora split into separate indexes so a large catalogue cannot poison concept queries; IDF-weighted coverage replaces term counting; a query term absent from the corpus raises the bar instead of lowering it. Calibrated on a hand-built gold set, then checked for robustness across the whole threshold range — a spike at one value is overfitting, not accuracy.

Method
RAG · IDF Weighting · Abstention · Recall / MRR Evals
Finding
Two-tier index (keyed + concept) prevents catalogue poisoning; IDF coverage > term counting; threshold robustness required
Reproduction
CyberRAG skill — /var/minis/skills/cyberrag/
Target
Calibrated abstention threshold stable across full gold set (no spike at single value)
ETA
2026-10-18
ACTIVE

Small-model routing and needle retrieval

Testing whether cheap specialist models can be routed in place of one general model — routing ablations against a single-model baseline, measured honestly rather than by a single shared prompt. Alongside it, adversarial long-context retrieval to find where a fine-tuned small model stops holding facts. Training runs on constrained hardware, so cost is a real constraint and not an abstraction.

Method
SFT · Routing · Adversarial Evals · On-Device Inference
Finding
Routing ablations needed; single-prompt evals are misleading; adversarial context finds the actual floor
Reproduction
MX local-model stack — shared/mx-local-model/
Target
Routing ablation framework + adversarial needle benchmark published
ETA
2026-10-22
ACTIVE

Wormlab — capability floor of small models

The gap. Public claims that AI-driven worms are feasible with small free models test a single model, deliberately withhold its size, and stop. Untested: the quantitative floor, the curve rather than the point, micro-policies against language models, and payload execution against rhetoric. A headline that cannot be reproduced is not a result — so the target is the sweep, not the demonstration.

P0 priority as of 2026-10-03: MX routing logic depends on this data.

Method
Model sweep: sizes [1B, 3B, 7B, 14B, 32B, 70B] × task classes [recon, exploit-dev, payload-gen, post-exfil, evasion] × 50 trials each. Gold set: 200 hand-labelled cyber tasks. Compute: single A100 80GB via Colab Pro.
Finding
Open question — no quantitative floor established in public literature
Target
Confusion matrix per (model_size, task_class) pair; publish sweep dataset + analysis
ETA
2026-10-25 (first results)
Reproduction
Experimental design doc — shared/wormlab/
ACTIVE

Critical infrastructure exposure census

Passive-first measurement of Algerian critical-sector exposure — energy, water, ports, rail, telecom edge — from public search indexes, DNS and historical records, certificate transparency, archived pages and published technical documentation. No active scanning of critical infrastructure, no authentication attempts. The output is a documented chain: observation → attribution → technology → known vulnerability → duration observable. Disclosure to the responsible party and the regulator before publication.

Method
Exposure · Reproducible Method · CERT-CC / ARPCE Route
Finding
Interim: gov.dz _dmarc NXDOMAIN; naftal.dz/baridi.dz/eldjazair.com.dz/mobilis.dz no DMARC; sonatrach.dz p=none; mf.gov.dz/poste.dz p=reject; djezzy.dz/ooredoo.dz p=quarantine; mdn.gov.dz SERVFAIL on TXT. Duration observable via Vercel edge age header on impersonation case demonstrates method.
Target
5-sector DMARC/SPF/DKIM census + duration observables for top 20 assets
ETA
2026-10-31
Reproduction
DoH queries via cloudflare-dns.com/dns-query + dns.google/resolve
COMPLETE

Duration and linkage on a live impersonation operation

A passive procedure for taking a suspected brand-impersonation operation from "this looks like phishing" to a documented record of what is linked, since when, and what its code actually does — without touching the operation, its users, or its data. Generator attribution, artifact correlation, the claim-implementation mismatch test, and duration inference, demonstrated on a live Algerian case. Boundaries stated, limits included, individual attribution left unestablished.

Method
Passive OSINT · Bundle Analysis · Deployment Correlation · Evidence Sequencing
Finding
v0-generated (generator meta); claim-implementation mismatch (no payment gateway exists); duration via Vercel edge age header
Reproduction
Method paper · Evidence manifest in workspace/mahata/
COMPLETE

Evidence grading for breach claims

How to tell a live disclosure from a re-upload. The standard applied to a public breach claim: trace the original actor and date independently, detect whether the circulating post is a repost of an earlier claim, confirm the named organisation exists and is genuinely affected, and state plainly what remains technically unverified. Confidence and provenance are recorded separately and never merged.

Method
OSINT · Verification · Provenance · Confidence
Finding
Wateen leak claim (Sep 2026) = re-upload of May 2026 claim; original actor lulzintel vs reposter updap unresolved
Reproduction
Evidence file — workspace/wateen-case/evidence.md
AUTHORIZED

Android attack-surface analysis

Exported components, cross-application data flows and authorization boundaries in mobile applications, verified with static and dynamic testing under written authorization only.

Method
AppSec · Authorized Engagements Only
Finding
Per-engagement — not published
Reproduction
android-ui-automation skill + Shizuku CLI

Handled directly. I do not sell access to findings, do not trade in credentials, and do not operate against infrastructure I am not authorized to test.

  • Contactsecurity@marwan-naili.me — machine-readable at /.well-known/security.txt
  • ScopeReports accepted in any infrastructure. Findings concerning Algerian critical infrastructure are routed through CERT-CC / ARPCE rather than published first.
  • TimingAcknowledgement within 3 business days. Publication after 90 days of vendor silence, unless a fix or mitigation exists.
  • MethodNon-intrusive by default — no exploitation, no authentication attempts, no denial of service, no state-affecting action.
  • PromiseNo legal action against researchers acting in good faith within this policy, and no attribution without consent.
Verify the target
A passing command proves the mechanism ran. The state after the action is the evidence. I test the objective, not the tool that was supposed to reach it.
Separate the layers
Observed, reported, inferred, unknown. Conflicts are preserved until resolved rather than overwritten by the newer, less complete account.
Prefer reversibility
Establish current state, preserve required artifacts, know the recovery path, then act. Recoverability is not traded away for convenience without a reason.
Environment is capability
Repeated friction is a missing primitive, not a reason to keep stacking hacks. The same failure twice is a bug in the approach.