Ingestion guide¶
How to point your SIEM / SOAR / pipeline at siemulator so it starts consuming synthetic alerts in real-vendor shapes — for integration testing, demo environments, training labs, or detection-engineering soak tests.
There's a live demo at
https://siemulator-y7uhf.ondigitalocean.app with default tokens
(logscale-dev-token, qradar-dev-token) — point your tool at it
before standing up your own instance to validate the integration end-
to-end without operational risk.
Synthetic data warning. Every alert is fabricated. Don't run alerts ingested from siemulator into a production detection pipeline, SIEM correlation rule, or analyst queue you can't reset. Pin the
X-Mock-Source: siemulatorheader /x-mock-sourcefield in every consumer as your "this is fake" guard — see Verification.
Table of contents¶
- Quickstart
- Choosing a surface
- Authentication patterns
- Polling patterns + dedup
- Platform recipes
- IBM QRadar SOAR / Resilient
- Splunk SOAR (Phantom)
- Cortex XSOAR (Demisto)
- Microsoft Sentinel — Logic App + custom logs
- Splunk Enterprise — REST modular input
- Elastic Stack — Logstash http_poller
- Tines / n8n / Zapier
- Custom Python poller
- Verification
- Going to production
Quickstart¶
If you just want to see whether your tool can talk to siemulator:
# These two endpoints are unauthenticated — your tool should hit them
# first to validate connectivity before negotiating credentials.
curl https://siemulator-y7uhf.ondigitalocean.app/logscale/api/v1/status
curl https://siemulator-y7uhf.ondigitalocean.app/qradar/api/help
If those work, your tool can reach siemulator. The recipes below configure each platform to poll the alert / offence endpoints on a schedule.
Choosing a surface¶
| You want to test | Use surface | Why |
|---|---|---|
| QRadar offence ingestion (most SOARs) | /qradar/* |
Native shape; consumers don't need custom mapping. 62 multi-source scenarios available with stable IDs for dedup. |
| Humio / Falcon LogScale push integrations | /logscale/* |
Mirrors Humio REST exactly — @timestamp, @id, @rawstring, #repo envelope. |
| CrowdStrike Falcon vendor pack | /alerts/entities/alerts/v2 |
Falcon Alerts v2 envelope ({meta{pagination,writes}, resources[], errors[]}) with DetectsAlert field names — composite_id identity, 0-100 int severity + severity_name, flat ATT&CK fields plus mitre_attack[], Falcon's filename/cmdline/sha256/hostname/user_name. /crowdstrike/api/v1/detects is kept as an alias. |
| Microsoft Defender vendor pack | /v1.0/security/alerts |
Graph Security v1.0 envelope ({@odata.context, @odata.count, value[]}) with the Alert entity's camelCase fields — vendorInformation.provider, Graph's severity/status enums, and the nested userStates[] / hostStates[] / networkConnections[] / processes[] / fileStates[] collections. /defender/api/security/v1.0/alerts is kept as an alias. |
| RSA NetWitness vendor pack | /rest/api/incidents |
NetWitness Respond envelope ({items[], pageNumber, pageSize, totalPages, totalItems, hasNext, hasPrevious}) with Respond incident field names — INC-<n> id, title/detail, riskScore+priority, entities under events[]. /netwitness/api/v1/incidents is kept as an alias. |
| Detection content (templates) — alert-shape testing | LogScale / QRadar | Both aggregator surfaces draw from the same 6-template pool; pick the shape your downstream consumer expects. |
| Multi-source attack-narrative testing (5-alert chains across multiple vendors) | /qradar/* with ?scenarios=… |
The QRadar surface exposes the whole 62 scenario library in one envelope, so cross-vendor correlation logic can be exercised end-to-end. |
Search-API testing (/queryjobs, Ariel) |
LogScale / QRadar | LogScale queryjobs POST→poll matches Humio; QRadar Ariel matches the IBM async-search shape. |
Range-paginated ingestion (Range: items=N-M) |
/qradar/* |
QRadar canonical pagination header. LogScale uses ?limit=N query param instead. |
Most SOAR integrations want /qradar/* because (a) QRadar's
offence shape is what most SOAR-vendor connectors are already built
against, and (b) the 62 multi-source scenarios let you test correlation
logic, not just shape parsing.
When to reach for the vendor-native endpoints: if your SOAR has a
vendor-specific parser (Falcon YAML pack, Defender Graph pack,
NetWitness decoder) that keys off the vendor's actual API envelope
shape, point that ingestion action at the vendor endpoint. The alert
payloads are the same — only the outer envelope changes to match
each vendor's real API. Your parser sees resources[] /
value[] / items[] in the shape it expects. All three
support ?scenarios=all|batch|replay with per-vendor rotation +
dedup state (independent of the QRadar counter). Reset a specific
vendor's state via POST /_debug/reset_vendor?vendor=crowdstrike
(or ?vendor=all).
Grading mode — don't let the corpus leak its own answers¶
Every curated scenario carries a ground-truth label. By default that
label is visible in the alert: the offence description begins with
a [SCENARIO-ID — 103 Ransomware — ...] tag, and that description
becomes the incident subject downstream. A consumer whose classifier
reads the subject can therefore score "correct" by echoing the label out
of its own input — which makes any category-accuracy number meaningless.
Use ?labels=blind on any run where you intend to measure
classification:
| Mode | Alert body | Use when |
|---|---|---|
| (default) | Full — scenario tag, category strings, _test_meta, expected_* all present |
Eyeballing the corpus; grading by grep on _test_meta |
?labels=strip |
Explicit grading metadata removed (_test_meta, test_notes, expected_*) |
You want the answer key out of the body but keep the scenario tag |
?labels=blind |
Also removes the implicit leaks: the [scenario — category] tag in the description, the category strings, the mock log-source name, and the correlation handles |
Measuring category / assessment accuracy |
The answer key is never lost — only moved. Every response carries an
X-Mock-Labels header: base64(JSON) keyed by offence id, holding the
full envelope (category_id, assessment, severity_band,
is_true_positive, must_extract_iocs, rationale). A grader joins it
on the offence id (_offense_id on the vendor-native shapes, which is
preserved in blind mode precisely so the join stays explicit):
curl -sD /tmp/h "$SIEM/qradar/api/siem/offenses?token=$TOK&scenarios=replay&labels=blind" -o /tmp/alerts.json
grep -i '^x-mock-labels:' /tmp/h | cut -d' ' -f2 | base64 -d | jq '."90131"'
# => {"category_id":103,"assessment":"MALICIOUS",...}
Blinding does not damage the alert: the vendor source, the narrative, the technical detail (CVEs, command lines, hashes) and all IOCs are preserved, so the consumer still has everything it needs to classify on the merits. Only the answer is removed.
Authentication patterns¶
siemulator accepts three auth channels per surface, and the same token works on either surface (cross-acceptance — convenience for mixed-vendor environments). Pick whichever your consumer ships natively:
| Channel | Header / location | Best for |
|---|---|---|
Authorization: Bearer <token> |
HTTP header | LogScale-shape consumers, REST clients with stock Bearer support |
SEC: <token> |
HTTP header | QRadar canonical — most existing QRadar consumers default to this |
?token=<token> |
URL query param | Proxies / SaaS connectors that strip Authorization / Sec-* headers in egress |
If your environment puts a forward proxy between your consumer and siemulator, default to the query-param channel — it survives every header-stripping proxy without code changes.
Cross-token acceptance means if your consumer was configured against
a legacy token from a previous integration and you accidentally paste
it as siemulator's SIEMULATOR_QRADAR_TOKEN, it still works.
This isn't laziness — it's a deliberate decision because both surfaces
serve synthetic data and forgiving config-paste mistakes removes a
class of friction during integration setup.
Polling patterns + dedup¶
The hardest part of any "poll a mock SIEM for alerts" integration
isn't the HTTP call — it's making sure your consumer doesn't create
12 incidents from one scenario when your cron runs every 60 seconds.
siemulator provides four modes via ?scenarios=… on
/qradar/api/siem/offenses to make this safe.
| Mode | Behaviour | When to use |
|---|---|---|
(default — no ?scenarios=) |
Returns N synthetic offences from the 6-template pool. Random IDs, random content per call. | Shape-only soak testing; load-style polling where you want a constant trickle. |
?scenarios=all |
One-shot dedup. Returns scenarios with offence IDs your process hasn't served yet. Each ID emitted once per process lifetime. Subsequent polls return [] until reset. |
Default for SOAR ingestion testing. A cron poll every 60s drains the 62 scenarios over ~62 polls, then quiesces. Your SOAR sees 62 distinct incidents, not the same one re-ingested 62 times. |
?scenarios=batch |
Round-robin — one scenario per call, rotating through the pool. | Slow-drip ingestion where you want a steady stream of fresh content. |
?scenarios=replay |
All 62 scenarios in one response, every call. | Bulk-load tests; one-shot end-to-end runs. |
?scenarios=mix |
All scenarios + N synthetic templates from Range: items=0-N. |
Mixed-pool stress testing. |
If you choose ?scenarios=all and want to replay the pool (e.g. after
adding new scenario variants, or as part of a CI reset), call:
The admin key is set via SIEMULATOR_ADMIN_KEY; the endpoint is 403
when unset.
Cron cadence recommendation: every 60 seconds is plenty.
siemulator handles thousands of req/s on basic-xxs, but your SOAR's
ingestion-cost ceiling typically matters more than siemulator's.
Platform recipes¶
IBM QRadar SOAR / Resilient¶
Resilient (and the standalone QRadar SIEM ingestion add-on) consumes the QRadar offence shape natively. Point it at siemulator like a real QRadar console:
| Setting | Value |
|---|---|
| Host | siemulator-y7uhf.ondigitalocean.app (or your instance) |
| Port | 443 |
| Auth | SEC header — set the SEC header to your SIEMULATOR_QRADAR_TOKEN |
| Verify TLS | Yes (DO App Platform serves a valid Let's Encrypt cert) |
| Polling endpoint | /qradar/api/siem/offenses?scenarios=all |
| Polling interval | 60 seconds |
What lands: every poll, your SOAR ingests fresh scenario offences as
they're drained from the pool. Each offence carries a _scenario_id
field — group incidents by that label to reconstruct multi-alert
narratives (S1 = 5 alerts, S5 = 4 alerts).
Splunk SOAR (Phantom)¶
Use the HTTP/REST asset type (built-in, no app install needed).
Asset configuration:
| Field | Value |
|---|---|
| Asset name | siemulator |
| Asset type | REST |
| Base URL | https://siemulator-y7uhf.ondigitalocean.app |
| Verify server certificate | true |
| Default headers | SEC: <your token> |
Playbook block (paste into a code block in a playbook):
import phantom.rules as phantom
import json, urllib.request
def poll_siemulator(action=None, success=None, container=None,
results=None, handle=None, filtered_artifacts=None,
filtered_results=None, custom_function=None, **kwargs):
url = "https://siemulator-y7uhf.ondigitalocean.app/qradar/api/siem/offenses?scenarios=all"
req = urllib.request.Request(url, headers={"SEC": "qradar-dev-token"})
with urllib.request.urlopen(req, timeout=10) as resp:
offences = json.loads(resp.read())
for o in offences:
# Pin the mock-source marker so you can never confuse this for
# production data downstream:
assert o.get("x-mock-source") == "siemulator", "real-SIEM source detected!"
cid = phantom.create_container(
label="events",
name=f"[{o['_scenario_id']}] {o['description'][:80]}",
description=o["description"],
severity={9: "high", 7: "medium", 5: "low"}.get(o["severity"], "low"),
artifacts=[{
"name": "offence",
"cef": {
"deviceCustomString1": o["_scenario_id"],
"deviceCustomString2": o["_detection"]["TechniqueId"],
"sourceAddress": o["source_ip"],
"destinationAddress": o["destination_ip"],
"severity": o["severity"],
"externalId": str(o["id"]),
},
"raw": o,
}],
)
return phantom.set_status(container=container, status="closed")
Schedule the playbook on a 1-minute cron.
Cortex XSOAR (Demisto)¶
Use the Generic Webhook / Generic REST integration (built-in).
Integration instance:
| Field | Value |
|---|---|
| Name | siemulator-incidents |
| Server URL | https://siemulator-y7uhf.ondigitalocean.app |
| Fetches incidents | true |
| Incident type | Alert |
| Mapper (incoming) | (see below) |
| First fetch | 1 minute ago |
| Fetch incidents (interval) | 1 minute |
Pre-process script (paste into the integration's script field):
import demistomock as demisto
import json
import requests
URL = "https://siemulator-y7uhf.ondigitalocean.app/qradar/api/siem/offenses"
HEADERS = {"SEC": demisto.params().get("apikey")}
def fetch_incidents():
last_run = demisto.getLastRun() or {}
resp = requests.get(f"{URL}?scenarios=all", headers=HEADERS, timeout=10)
resp.raise_for_status()
offences = resp.json()
incidents = []
for o in offences:
assert o["x-mock-source"] == "siemulator"
incidents.append({
"name": f"[{o['_scenario_id']}] {o['description'][:80]}",
"occurred": demisto.toIso8601(o["start_time"] / 1000),
"rawJSON": json.dumps(o),
"severity": {9: 3, 7: 2, 5: 1}.get(o["severity"], 1),
"type": "Alert",
"details": o["description"],
"labels": [
{"type": "scenario_id", "value": o["_scenario_id"]},
{"type": "technique_id", "value": o["_detection"]["TechniqueId"]},
],
})
demisto.incidents(incidents)
demisto.setLastRun({"last_id": offences[-1]["id"] if offences else last_run.get("last_id")})
fetch_incidents()
The one-shot ?scenarios=all mode means after 38 polls the pool
drains and fetch_incidents stops creating incidents until you reset
the served-set via /qradar/_debug/reset_scenarios. Useful — your CI
test sees exactly 22 incidents, then quiesces.
Microsoft Sentinel — Logic App + custom logs¶
Use a scheduled Logic App to poll siemulator and forward to a Sentinel custom log table.
Logic App workflow (Designer or JSON):
- Trigger: Recurrence every 1 minute.
- Action: HTTP —
- Method:
GET - URI:
https://siemulator-y7uhf.ondigitalocean.app/qradar/api/siem/offenses?scenarios=all - Headers:
SEC=qradar-dev-token - Action: Parse JSON — schema from the QRadar response (the easiest way is to run once, copy the body from the run history, and let the designer generate the schema).
- Action: For each on
body('Parse_JSON')→ Send Data (Azure Log Analytics Data Collector) — - Workspace ID + key from your Sentinel workspace
- Custom log name:
siemulator_offence_CL - JSON request body:
items('For_each')
Sentinel then exposes the data as siemulator_offence_CL in KQL:
If xmocksource_s ever returns anything other than "siemulator"
in production, alert-on-it — that's the leak indicator.
Splunk Enterprise — REST modular input¶
Use the Splunk Add-on Builder or the built-in REST API Modular Input add-on (Splunkbase, free) to poll siemulator.
Input configuration:
| Field | Value |
|---|---|
| Endpoint URL | https://siemulator-y7uhf.ondigitalocean.app/qradar/api/siem/offenses?scenarios=all |
| HTTP method | GET |
| Authentication | none (we use a custom header instead) |
| Custom headers | SEC=qradar-dev-token |
| Interval | 60 (seconds) |
| Response handler | default (JSON array) |
| Source type | siemulator:offence |
| Index | mock_security (separate from your prod indexes) |
In Splunk search:
index=mock_security sourcetype=siemulator:offence
| stats count by _scenario_id, _detection.TechniqueId
| sort -count
Elastic Stack — Logstash http_poller¶
# /etc/logstash/conf.d/siemulator.conf
input {
http_poller {
urls => {
siemulator => {
method => get
url => "https://siemulator-y7uhf.ondigitalocean.app/qradar/api/siem/offenses?scenarios=all"
headers => {
SEC => "qradar-dev-token"
}
}
}
request_timeout => 10
schedule => { every => "60s" }
codec => "json"
tags => ["siemulator", "synthetic"]
}
}
filter {
# Hard fail if anything reaches here without the mock-source marker.
if ![x-mock-source] or [x-mock-source] != "siemulator" {
drop {}
}
mutate {
rename => { "[_detection][TechniqueId]" => "mitre_technique_id" }
rename => { "[_scenario_id]" => "scenario_id" }
}
}
output {
elasticsearch {
hosts => ["https://your-es:9200"]
index => "mock-siemulator-%{+YYYY.MM.dd}"
user => "logstash_writer"
password => "${ES_LOGSTASH_PWD}"
}
}
Kibana Discover query: tags:siemulator AND scenario_id:S*.
Tines / n8n / Zapier¶
Visual workflow tools — pattern is identical across all three:
- Trigger: Schedule every 60 seconds.
- Action — HTTP Request:
- URL:
https://siemulator-y7uhf.ondigitalocean.app/qradar/api/siem/offenses?scenarios=all - Method:
GET - Headers:
SEC: qradar-dev-token - Action — Loop over items (array body).
- Action — Branch by
_scenario_idor push each offence to your downstream tool (Slack alert, ticket-create, webhook to another workflow).
In Tines specifically, the ?scenarios=all one-shot dedup means your
storyboard runs end-to-end with exactly 22 incidents, then quiesces —
useful for iterating on a workflow without re-creating 22 Jira
tickets every minute.
Custom Python poller¶
Minimal reference — drop into any container, runs forever, prints one line per ingested offence:
"""Minimal siemulator → stdout poller. ~40 LOC, stdlib only.
Run with:
python siemulator_poll.py https://siemulator-y7uhf.ondigitalocean.app qradar-dev-token
"""
from __future__ import annotations
import json, sys, time
import urllib.request, urllib.error
def poll(base_url: str, token: str, interval_s: int = 60) -> None:
url = f"{base_url.rstrip('/')}/qradar/api/siem/offenses?scenarios=all"
req = urllib.request.Request(url, headers={"SEC": token})
seen: set[int] = set()
while True:
try:
with urllib.request.urlopen(req, timeout=10) as resp:
offences = json.loads(resp.read())
except urllib.error.HTTPError as e:
print(f"[{time.strftime('%H:%M:%S')}] poll failed: HTTP {e.code}", flush=True)
time.sleep(interval_s)
continue
for o in offences:
assert o.get("x-mock-source") == "siemulator", "non-mock data!"
if o["id"] in seen:
continue
seen.add(o["id"])
print(
f"[{time.strftime('%H:%M:%S')}] "
f"id={o['id']} scenario={o['_scenario_id']} "
f"sev={o['severity']} desc={o['description'][:80]}",
flush=True,
)
time.sleep(interval_s)
if __name__ == "__main__":
poll(sys.argv[1], sys.argv[2])
After 38 polls (~38 minutes), the pool drains; call
/qradar/_debug/reset_scenarios with the admin key if you want to
replay the whole library.
Verification¶
After wiring up any of the recipes above, verify ingestion from both sides:
Siemulator side — quick triage (QRadar only, last 100 requests):
curl -H "X-Admin-Key: $ADMIN_KEY" \
https://your-siemulator/qradar/_debug/recent \
| jq '.requests[:5] | .[] | {ts, path, mode, response_count, client}'
Siemulator side — full access log across both surfaces (5000 recent requests + aggregations):
# Who is consuming what?
curl -fsS -H "X-Admin-Key: $ADMIN_KEY" \
https://your-siemulator/api/access-log/stats \
| jq '{ total, top_clients, top_user_agents, by_auth, by_status,
duration_ms }'
# Show only YOUR consumer's traffic (filter by IP)
curl -fsS -H "X-Admin-Key: $ADMIN_KEY" \
"https://your-siemulator/api/access-log?limit=50" \
| jq '.entries[] | select(.client_ip == "203.0.113.45")'
# Show only failed requests (helps debug "my consumer can't auth")
curl -fsS -H "X-Admin-Key: $ADMIN_KEY" \
"https://your-siemulator/api/access-log?status=401&limit=10" | jq .
The access log records timestamp, path, redacted query, auth channel
name (bearer / sec / query / none — never the token value),
client IP (X-Forwarded-For aware), user-agent, status, duration in
ms, and response bytes. If your consumer's IP / user-agent doesn't
show up in top_clients / top_user_agents, the request isn't
reaching siemulator (firewall? wrong URL? DNS?).
The same records also go to stdout as structured JSON lines —
docker logs siemulator | jq . or doctl apps logs <app-id>
--follow will tail them in real time, which is often the fastest
way to debug "why isn't my poller hitting siemulator?" during
initial integration setup.
Consumer side — confirm your tool can read the mock-source marker:
| Tool | Query |
|---|---|
| Splunk | sourcetype=siemulator:offence "x-mock-source"="siemulator" \| stats count |
| Sentinel | siemulator_offence_CL \| where xmocksource_s == "siemulator" \| count |
| Elastic | tags:siemulator AND x-mock-source:siemulator |
| XSOAR | filter incident by labels.scenario_id:S* |
If x-mock-source is missing or != siemulator, stop the
ingestion — something has either pointed your consumer at a real
SIEM by mistake, or rewritten the response in transit.
The pinning recommendation is to fail-closed: if a single ingested record arrives without the marker, page the ingestion owner.
Going to production¶
When you graduate from "exploring with the live demo" to "depending on siemulator in CI / staging / a long-running test environment":
-
Stand up your own instance. Don't depend on the public demo URL for anything load-bearing. Deploy options are in the README § Deploy on DigitalOcean App Platform; the Docker image is at
ghcr.io/sirp-labs/siemulator:latest. -
Rotate the tokens. The public demo uses
logscale-dev-token/qradar-dev-token— public sentinels. SetSIEMULATOR_LOGSCALE_TOKENandSIEMULATOR_QRADAR_TOKENto fresh values (openssl rand -hex 24) on your own deploy. -
Set
SIEMULATOR_ADMIN_KEY. Without it, the/qradar/_debug/*endpoints return 403 (safe default). Set it if you want request-capture for ingestion debugging or scenario-set reset capability. -
Front it with HTTPS. DO Apps does this automatically; on your own infra, terminate TLS upstream.
-
Pin the
x-mock-sourcemarker in every consumer's parser, with a fail-closed branch if the field is missing or !=siemulator. This is your "the mock isn't accidentally proxying real data" guard. -
Don't multi-instance unless you don't use
?scenarios=all. The one-shot dedup state is per-process; running 2 instances behind a load balancer means each instance serves each scenario independently (your consumer sees the same offence twice). Stick toinstance_count: 1for the one-shot mode; bump for shape-only soak testing where every poll returning random offences is fine. -
Snapshot-pin in CI. Once your integration parses siemulator responses correctly, lock the parse shape with a contract test that fetches
/qradar/api/siem/offenses?scenarios=replayand asserts on the field shape — that way if siemulator ever changes its output in a way that breaks your consumer, you find out in your test suite (not in production).