ServiceNow
Incidents and change records kept current, recent changes as evidence, knowledge articles in every diagnosis, fixes ordered through the catalogue, and ServiceNow's own failures handled.
ServiceNow is where a bank records what happens to its systems, so 3AM works through it:
- The incident record. 3AM attaches to the open incident on the service if a pager or monitoring integration already opened one. Otherwise it opens one. Every step goes in as a work note: evidence, proposal, approval, action, verification.
- What changed. Every incident shows the changes on the service and the CIs it depends on in the last 24 hours, with the running ones first. The first question in an outage is "what changed?", and the approver sees the answer.
- Change records for 3AM's own actions. Where your policy asks for one, 3AM opens a standard change from a pre-approved template, or an emergency change, before anyone approves. A freeze or blackout window shows up then, not afterwards.
- Your knowledge base. Published articles, such as runbooks and known errors, are searched in every diagnosis, and the incident note cites them.
- Fixes that can only be requested. Some fixes, like group membership, an account unlock, a certificate or a firewall rule, go through the service catalogue. 3AM orders the item, ServiceNow runs its own approval and fulfilment, and 3AM waits for the result.
- ServiceNow itself. 3AM notices when ServiceNow stops answering and proves why. It keeps every ticket update until ServiceNow is back. A MID Server that goes down is restarted, and stuck request fulfilment is traced to its cause.
Before you start
- A ServiceNow admin, once, to create the integration user and an OAuth client.
- Outbound HTTPS from 3AM to your instance. ServiceNow never connects to 3AM.
- For MID Servers that run in Kubernetes, the Kubernetes connector with the MID Server's namespace in its namespaces, so 3AM can restart them.
1. Create the integration user
In ServiceNow, create a user for 3AM (User Administration → Users). Tick Web service access only and give it the roles for the features you'll use:
| Feature | Role |
|---|---|
| Incidents, changes, problems, CMDB | itil |
| Knowledge bases for SONE | knowledge (or user criteria that let it read the bases you name) |
| Request fulfilment: requests and flow runs | itil and flow_operator |
| Ordering catalogue items for fixes | The items' own user criteria |
| MID Server status | Read on ecc_agent |
Sign in with OAuth client credentials (recommended):
- System OAuth → Application Registry → New → Create an OAuth API endpoint for external clients. Name it
3AMand keep the client secret. - Set OAuth Application User to the integration user. Client-credential tokens act as this user.
- Make sure the system property
glide.oauth.inbound.client.credential.grant_type.enabledistrue.
A client certificate (mutual TLS, mapped to the user) or a password also work. Passwords expire and get locked, so prefer OAuth.
If 3AM will open change records, create a standard change template for its fixes (Change → Standard Change →
Template proposal), named for example 3AM remediation, and have it approved like any other template.
2. Add the connector
In the console: Settings → Add a connection → ServiceNow. Or in connectors.json:
{ "name": "servicenow", "kind": "servicenow",
"settings": { "url": "https://bank.service-now.com",
"client_id": "f1c2…",
"client_secret": { "file": "/etc/3am/secrets/servicenow-client-secret" },
"assignment_group": "Payments Ops",
"standard_change_template": "3AM remediation",
"knowledge_bases": ["IT Runbooks"],
"watch_fulfilment": true,
"catalog_items": ["Add user to AD group", "Unlock service account"] } }catalog_items is an allow-list: 3AM never orders anything else.
kind: servicenowticketschangeCMDBcodealerts inactions| Setting | What to enter | Required | Default |
|---|---|---|---|
url | Instance URL, e.g. https://bank.service-now.com | yes | — |
client_id | OAuth client id (Application Registry) | no | — |
client_secretsecret | OAuth client secret | no | — |
username | Integration user (password sign-in, or the OAuth password grant) | no | — |
passwordsecret | Password | no | — |
cert_file | Client certificate for mutual TLS (PEM) | no | — |
key_file | Its private key (PEM) | no | — |
ca_file | CA for a TLS-intercepting proxy (PEM) | no | — |
assignment_group | Assignment group for incidents 3AM opens (name or sys_id) | no | — |
attach_window_h | Attach to an open incident on the service opened in the last N hours | yes | 24 |
resolve_on_verified | Resolve the incident after a verified fix (else: a work note saying it is ready to resolve) | yes | false |
resolve_code | Resolution code used when resolving | yes | "Solution provided" |
change_model | Change records for 3AM's actions (when policy asks): standard or emergency | yes | "standard" |
standard_change_template | Pre-approved standard change template (name or sys_id) | no | — |
service_tables | CMDB tables that hold services | yes | ["cmdb_ci_service","cmdb_ci_service_discovered"] |
knowledge_bases | Knowledge bases to index for SONE (titles; empty: none) | no | — |
mid_servers | MID Servers to watch (names; empty: all) | no | — |
watch_fulfilment | Watch request fulfilment (stuck requests, failing flows) | yes | false |
catalog_items | Catalogue items 3AM may order for a fix (names or sys_ids) | no | — |
event_alerts | Take Event Management alerts (em_alert) as signals | yes | false |
unreachable_after_s | Seconds failing before ServiceNow itself is an incident | yes | 120 |
mid_stale_after_s | Seconds without a MID Server heartbeat before it is down | yes | 300 |
request_stuck_after_s | Seconds an approved request may wait for fulfilment | yes | 3600 |
stuck_requests | Stuck requests (of one item) that make an incident | yes | 3 |
flow_errors | Failed flow runs in an hour that make an incident | yes | 3 |
fulfil_wait_s | How long order_item waits for fulfilment | yes | 900 |
Change records
Whether 3AM opens a change record is policy, per service or per action:
"change_records": { "default": "none", "freeze": "warn" },
"rules": [ { "match": { "environment": "prod" }, "change": "standard" } ]When a change is wanted, 3AM:
- Opens the change before asking for approval. It describes the proven cause, the evidence, the action, the rehearsal, and how 3AM will verify the fix.
- Runs ServiceNow's conflict check. A freeze or blackout conflict is shown in the approval request
(
"freeze": "warn"), or stops 3AM altogether ("freeze": "block"). - Moves the change to Implement once approved, and only then acts. If the change can't move, nothing runs.
- Closes the change as successful or unsuccessful, with the verification result. If the fix is rejected or blocked, the change is cancelled.
If policy wants a change and none can be opened, 3AM doesn't act.
3. Verify
ServiceNow Zurich, signed in with oauth; incidents readableTest connection also reads every table the features you turned on need. A missing role shows up there, not at 3am during an incident.
What it detects
| Signal | Meaning |
|---|---|
ServiceNowUnreachable | ServiceNow hasn't answered for unreachable_after_s (2 min), with the cause: DNS, connection, TLS, timeout, server errors, or a hibernating instance |
ServiceNowAuthFailing | ServiceNow refuses 3AM's credentials |
ServiceNowPermissionDenied | 3AM can sign in, but a role is missing for a table it needs |
ServiceNowRateLimited | An inbound REST rate limit is throttling 3AM |
ServiceNowMIDServerDown | A MID Server is Down, or hasn't sent a heartbeat for mid_stale_after_s (5 min) |
ServiceNowRequestsStuck | Approved requests for one item have waited longer than request_stuck_after_s (1 h) for fulfilment |
ServiceNowFlowErrors | Flow runs keep failing (flow_errors in an hour) |
| Event Management alerts | With event_alerts on, open critical and major alerts become signals, and 3AM's diagnosis goes back on the alert |
What it proves and fixes
| Cause | How 3AM knows | Fix |
|---|---|---|
| MID Server down | ServiceNow shows it Down, or its heartbeat is old | Restart its pod or workload (Kubernetes connector), then wait until ServiceNow shows it Up. On a VM: the restart command for its host |
| Requests stuck behind a down MID Server | Approved requests waiting, and the MID Server their flows use is down | Restart the MID Server, then check the requests move |
| Fulfilment credential refused | Flow runs fail with 401 or credential errors while the MID Server is up | Advice: renew the credential the flow's connection uses, then retry the requests |
| Fulfilment flow error | Flow runs fail for another reason | Advice: the flow's error, for its owner |
| ServiceNow unreachable | It hasn't answered for minutes, and the failure says why | Advice: whose problem it is (network, ServiceNow, or a hibernating instance). Ticket updates wait in 3AM |
| Credentials refused | 401 for minutes | Advice: what a ServiceNow admin checks (locked, expired, OAuth client) |
| Role missing | 403 on a table it needs | Advice: which role |
| Rate limited | ServiceNow answers 429 | Advice: the rate limit rule that's throttling 3AM |
When ServiceNow is down
ServiceNow is a hosted service, so 3AM can't restart it. Incidents still get handled:
- Diagnosis and fixes carry on. Approvals go through your other channels.
- Ticket updates are kept in order in 3AM's data directory: the new incident, each work note, the resolution. They're sent once ServiceNow answers again.
- Incidents opened during the outage show
pending:<episode>until ServiceNow gives them a number. The incident record is then updated. - A request ServiceNow rejects as wrong (for example a field that doesn't exist) isn't kept, because it would never succeed. It's in the audit log.
Actions
| Action | Parameters | Notes |
|---|---|---|
order_item | item, variables, note | Only items in catalog_items. The dry run checks the item and its variables. Waits up to fulfil_wait_s (15 min) for fulfilment, and reports a rejection in ServiceNow's approval as a failure |
instance_status, mid_status, fulfilment_status, ci_status, changes, incident_history, on_call, kb_search | Read-only evidence |
Certification
Live certification against ServiceNow Zurich is in progress. Until it finishes, this connector is tested against a model of ServiceNow's REST APIs, not a live instance.
Troubleshooting
| You see | Do this |
|---|---|
the instance refuses 3AM's credentials (oauth) | Check the Application Registry entry is active, its secret is the one 3AM has, and its OAuth Application User is set and active |
can't read: sc_req_item, sys_flow_context | Give the integration user itil and flow_operator, or turn watch_fulfilment off |
the instance is hibernating | A developer instance sleeps when unused: wake it from the developer portal |
no single active standard change template named … | Check standard_change_template matches the template's name exactly, or use its sys_id |
Incidents show pending:… | ServiceNow was unreachable when the incident started. 3AM sends the update when it answers |