When IP Lookups Fail: An Edge Policy Incident Runbook
By IP Geolocation Team · · 9 min read
Your application can keep serving requests while an IP lookup dependency fails. That sounds reassuring until a regional restriction disappears, every login gets challenged, or a campaign report labels unknown traffic as domestic. Availability alone does not tell you whether the application is making acceptable decisions.
Good IP geolocation API error handling preserves the difference between missing evidence and reassuring evidence. This runbook gives platform, fraud, product, and analytics teams a shared response to timeouts, incomplete results, and recovery. It describes custom API-based middleware for an edge worker, serverless service, or gateway you operate.
Start with evidence states, before allow or deny
Keep lookup status separate from the application's decision. A successful HTTP response may contain insufficient data for the rule you want to apply. Conversely, a failed enrichment call does not mean the visitor has committed abuse. Preserve both distinctions in the incident dashboard and policy code.
| Observed result | Evidence state | First action |
|---|---|---|
| Timeout or network failure | Unavailable | Apply the route fallback; inspect dependency reachability |
| Authentication/authorization rejection | Unavailable | Check secret deployment and account access; do not retry blindly |
| Throttling or quota rejection | Unavailable | Reduce calls and inspect usage and response details |
| Malformed JSON or wrong response shape | Invalid response | Inspect contract drift or an intermediary error page |
| Valid response without country | Location incomplete | Use the missing-location policy |
| No documented anonymity fields | Screening unknown | Never convert missing evidence into a clean result |
These are application classifications, not a promise about which HTTP status every provider returns. Consult the actual response and current API reference. Record an error category without copying API keys, credentials, or raw customer payloads into logs.
Choose degraded behavior for each business action
A language suggestion and a restricted download cannot share one fallback. Product owners should approve optional-personalization behavior. Security and regional-policy owners should approve access boundaries. Make these decisions before an on-call engineer has to trade customer access against an unknown dependency state.
| Workflow | When required evidence is missing | Preserve |
|---|---|---|
| Language or currency suggestion | Use saved preference or neutral selector | User choice and valid product pricing |
| Sensitive login | Use approved alternate verification | Authentication, rate limits and recovery safeguards |
| Checkout risk review | Use the established review or hold path | Payment controls and duplicate-submission protection |
| Region-restricted content | Withhold protected content pending sufficient evidence | The regional restriction and a clear retry path |
| Analytics enrichment | Mark unknown and queue permitted enrichment | Original event time and source provenance |
Those actions are design examples, not IP-Info.app product automations. A geography requirement also survives successful MFA: authentication establishes account control, not physical presence. Do not allow a login challenge to bypass an independent geographic access rule.
A five-step dependency incident workflow
1. Establish scope
Compare errors by deployment, region, route, and credential version. Separate upstream HTTP failures from your own deadlines and parser errors. Check whether a recent release changed the trusted IP source or the secret binding. A local configuration mistake should not trigger an unrelated vendor migration.
2. Contain the failing call path
Stop automatic retry amplification and pause nonessential backfills. A circuit breaker can temporarily stop calls to a failing dependency, but its scope matters. One worker instance's state is not necessarily shared across the deployment. Keep the breaker mechanism separate from the approved access policy.
3. Apply the route's approved fallback
Route incomplete location and unavailable lookups explicitly. Show a helpful retry or verification message for affected users. For checkout, use an existing review state rather than resubmitting a payment request. For optional localization, a saved preference often avoids changing the user experience at all.
4. Retain enough evidence to explain impact
Record the lookup state, failure category, route class, policy version, and chosen action. Measure affected requests and customer consequences separately. The edge decision observability guide covers the broader logging design; the incident record needs to answer which decisions changed and who must review them.
5. Restore service through a recovery gate
Verify credentials, contract shape, and representative responses before restoring normal volume. Recover a small approved portion of traffic first, then inspect errors and decision outcomes. One successful lookup against a public resolver does not prove recovery across customer regions or address families.
Implement a bounded lookup in trusted code
IP-Info.app documents this cURL request. The optional timing parameter asks for response metadata; it does not guarantee a response deadline. The public endpoint accepts IPv4 and IPv6. Keep the visitor IP explicit because an omitted address resolves the caller instead.
curl --fail-with-body --get \
'https://api.ip-info.app/v1-get-ip-details' \
--data-urlencode 'ip=8.8.8.8' \
--data-urlencode 'getPerformanceData=true' \
-H 'accept: application/json' \
-H "x-api-key: $IP_INFO_API_KEY"A TypeScript function can return the evidence state without choosing whether to admit the user. This example projects country only from the documented response. Its caller must validate the client IP, supply a secret and a measured deadline, then handle every state with the route policy.
type FailureCause = 'configuration' | 'http' | 'timeout' |
'transport' | 'invalid_response';
type CountryLookup =
| { state: 'available'; countryCode: string }
| { state: 'incomplete' }
| { state: 'unavailable'; cause: FailureCause; status?: number };
export async function lookupCountry(
clientIp: string, apiKey: string, deadlineMs: number,
): Promise<CountryLookup> {
if (!apiKey || !clientIp || !Number.isInteger(deadlineMs) || deadlineMs <= 0) {
return { state: 'unavailable', cause: 'configuration' };
}
const signal = AbortSignal.timeout(deadlineMs);
try {
const url = new URL('https://api.ip-info.app/v1-get-ip-details');
url.searchParams.set('ip', clientIp);
const response = await fetch(url, {
headers: { 'x-api-key': apiKey, accept: 'application/json' },
signal,
});
if (!response.ok) {
return { state: 'unavailable', cause: 'http', status: response.status };
}
const data: unknown = await response.json();
if (!data || typeof data !== 'object' ||
!('ip' in data) || typeof data.ip !== 'string') {
return { state: 'unavailable', cause: 'invalid_response' };
}
if (!('countryCode' in data) || data.countryCode == null) {
return { state: 'incomplete' };
}
if (typeof data.countryCode !== 'string' || !/^[A-Z]{2}$/.test(data.countryCode)) {
return { state: 'unavailable', cause: 'invalid_response' };
}
return { state: 'available', countryCode: data.countryCode };
} catch (error) {
return {
state: 'unavailable',
cause: signal.aborted ? 'timeout' :
error instanceof SyntaxError ? 'invalid_response' : 'transport',
};
}
}The full public response schema requires only ip and makes country optional. Therefore, incomplete is an expected state to support. An available country means location evidence exists; it does not mean VPN screening passed or the visitor is eligible for access.
The function performs one request and returns a failure category. Have the caller emit that category to your monitoring system and execute the approved fallback. Do not ignore the result or add an unconditional origin request after it. It is a lookup boundary, not a complete security middleware package.
Protect the trust boundary and the request budget
Accept the visitor address only from infrastructure you control and have configured to overwrite untrusted forwarding headers. Prevent direct origin access from bypassing that boundary. Otherwise, an attacker can choose the address you evaluate. A reliable lookup cannot repair an untrusted input.
Allocate the lookup deadline inside the route's overall response budget, including other dependencies. Measure end-to-end duration from the deployed runtime. Provider timing metadata is useful context but may exclude your network path, scheduling, and downstream work. No fixed timeout suits every login, download, and batch job.
Treat retries as additional requests that consume time and may incur usage. Inspect authentication and quota failures instead of immediately repeating them. If an asynchronous enrichment job can retry, use bounded attempts and backoff that respects server guidance. Never retry a checkout transaction just because its IP lookup failed.
Cloudflare's exception passthrough documentation describes forwarding to origin after an exception, with limits around consumed request bodies. That mechanism does not define your access policy. A custom gateway or edge implementation must enforce the same restriction when its enrichment dependency is unavailable.
Do not let stale or absent screening masquerade as certainty
Caching can reduce repeated calls, but a cached observation has an age and a purpose. Store when it was obtained and which source produced it. Do not reset freshness merely because you served it from cache. Approve separate freshness limits for harmless personalization and sensitive access decisions.
IP-Info.app's public marketing describes VPN, proxy, Tor, and risk-related workflows. Its curated lookup schema does not document matching detection flags or a numeric risk score. If your policy requires anonymous IP detection, verify the actual screening contract; this country-only example cannot satisfy that dependency.
Even an affirmative VPN result is not proof of fraud. Corporate gateways, privacy services, and shared networks can serve legitimate users. Combine supported screening evidence with your existing account and transaction controls. Missing screening remains unknown, regardless of whether the country is permitted.
For regional rules, define what evidence qualifies and what happens when it is absent. Network location is not GPS, legal residence, or proof of compliance. The regional access policy guide explains the broader access design; this runbook preserves its restrictions during a dependency incident.
Recover analytics without rewriting the incident
Keep unknown geography visible in business intelligence reports. Otherwise, an outage can look like a regional demand change. Preserve event time separately from later enrichment time, and label backfilled records. A lookup performed tomorrow is not proof of the network attributes present during yesterday's event.
Include ASN, organization, and ISP-related network analysis when reviewing suspicious traffic clusters. Those attributes provide context, not a user identifier. For ad fraud filtering, keep unevaluated events separate from accepted traffic so a screening outage does not improve reported quality by accident.
IP-Info.app markets bulk enrichment workflows, but its public reference does not specify a batch endpoint. Confirm the supported interface before planning recovery throughput. Any bounded job using the single-address API is custom processing, with its own queue and retry accounting.
Rehearse the runbook before the next incident
Test a delayed response, failed authorization, throttling, malformed JSON, missing country, and unavailable screening. Add stale cached evidence and a deployment with no secret. Check both IPv4 and IPv6 input handling. These are injected test cases, not claims that the provider exhibits each failure.
Assert the customer consequence: a protected response remains unavailable, a language selector still works, and a checkout does not execute twice. Use shadow-mode policy testing for observing proposed changes, then separately test enforced failure behavior. Record who can disable a rule and which restrictions must remain active.
Close the incident after queued work is reconciled and normal decision behavior returns, not merely when HTTP errors stop. If a fallback provider is involved, use the provider migration checklist to verify field meanings. A second source can introduce different country assignments and screening semantics precisely when operators have the least time to investigate.
Questions from the on-call handover
Should every route fail closed?
No single setting fits every action. Optional personalization can fall back to user choice. Protected content must retain its access boundary. Document each route's approved behavior and keep existing authentication, payment, and regional-policy checks independent of lookup availability.
Does this require a native edge integration?
No. This is a custom API-based pattern. The verified product surface is the HTTP lookup. Your team owns the worker, gateway, serverless handler, WAF coordination, secrets, monitoring, and recovery behavior; the example does not install a native Cloudflare or CDN connector.
Can we replace missing location with a default country?
Use a neutral experience for optional content, but preserve unknown location in evidence and analytics. Assigning a default country would manufacture a fact that downstream access rules and reports might trust. Let the user express a preference without relabeling it as a network observation.