SIP TROUBLESHOOTING TECHNICAL

SIP 503: Service Unavailable, Explained

SIPNEX ·

SIP 503 Service Unavailable means the server cannot handle the request for now — because of overload, maintenance, or exhausted capacity. RFC 3261 says the server MAY include a Retry-After header indicating when to try again, and the client should try an alternate server. On dialer trunks, a 503 usually means the carrier is saying “no room right now.”

It is the most retry-friendly failure in SIP. It is also the one most often misread in Asterisk logs. If you dial at volume, 503 is the code you see most on your worst days. This guide covers what it means, the cause ladder behind it, and the Asterisk mixup that sends teams chasing the wrong fix. For the class-by-class map of every code, start with the SIP response codes guide.

What RFC 3261 actually says about 503

RFC 3261 §21.5.4 defines 503 as a temporary condition: “The server is temporarily unable to process the request due to a temporary overloading or maintenance of the server.” Three details matter in practice:

  • It is temporary by definition. A 403 Forbidden is different: there, the RFC says authorization will not help. A 503 invites another attempt. Just not right away, and not always on the same route.
  • Retry-After is optional. The server MAY attach a Retry-After header. On a 503, §20.33 says it indicates “how long the service is expected to be unavailable to the requesting client.” When present, it is the most useful byte in the response. Without it, the RFC tells the client to treat the 503 like a 500. Carriers are not required to send it, so do not count on it.
  • The client should try an alternate server. The RFC says a client receiving a 503 “SHOULD attempt to forward the request to an alternate server.” That makes 503 the RFC’s built-in failover trigger. The spec also lets a server simply refuse the connection or discard the request rather than send a 503 at all. That is why capacity events can look like timeouts rather than clean rejections.

That is the protocol. What a carrier actually means by a 503 is convention. The conventions are steady enough to build a ladder from.

The 503 cause ladder for dialer operators

Carriers use 503 for a small set of cases. Telnyx’s published SIP response code list, for example, returns 503 when no route is available, when a calls-per-second limit is reached, and when termination fails further downstream. Exact meanings vary by carrier. Work the ladder top-down:

  1. CPS ceiling hit. Most wholesale trunks enforce a calls-per-second cap. A predictive dialer ramping a fresh list can burst far above its allowed rate. On many trunks, every call over the ceiling comes back 503. The tell: failures cluster at the start of dialing waves and vanish when pacing drops.

  2. Channel cap reached. On trunks that cap concurrency, the limit tracks live calls rather than dial rate. Long calls push you into this ceiling even at modest CPS. Carriers differ on the code here: some answer an over-limit call with 503, while Telnyx’s list returns 403 when an account is over its concurrent-call limit.

  3. Carrier-side congestion or maintenance. This is the textbook RFC case: the carrier’s own capacity is the limit. These events are global. Every destination fails at once, and they clear on their own.

  4. Upstream route failure surfaced as 503. When a carrier’s downstream partner fails to complete the call, some carriers, Telnyx among them, return the failure as a 503. The tell: failures cluster on certain destinations while the rest of your traffic completes.

  5. Account-level holds. A carrier can also fail every call while a billing hold or admin throttle is in place. Whether that arrives as a 503 or a 403 is carrier-specific, not an RFC meaning, so ask your carrier which code it uses. It is the one rung that never clears by waiting.

Note what is not on this ladder: authentication failures, IP allowlist misses, and caller-ID blocks. Carriers signal those with 403. That is a different code with a different playbook, covered in the SIP 403 Forbidden guide. Analytics-based call blocking has its own required code: SIP 603+ “Network Blocked”.

Asterisk AUTO-CONGEST is not a 503

This mixup is easy to make, and it sends troubleshooting the wrong way. Asterisk’s AUTO-CONGEST and a carrier 503 both end the call with a CONGESTION result. But they are different mechanisms with different fixes.

AUTO-CONGEST belongs to chan_sip, the older SIP channel driver. Asterisk removed chan_sip in version 21, but it still ships in Asterisk 20 and earlier, and VICIdial servers on those versions commonly still run it. PJSIP handles a silent far end differently, as covered below.

AUTO-CONGEST is a chan_sip no-response timer. When chan_sip sends an INVITE, it arms an auto-congest timer set to SIP timer B. That is 64×T1, or 32 seconds at the default 500 ms T1. It fires only if no response of any kind comes back before it expires. Any arriving response cancels it. An auto-congested call means your INVITE went into a void.

A real 503 takes a different path. When a 503 response actually arrives, Asterisk’s response handler maps it through its hangup-cause table to AST_CAUSE_CONGESTION (ISDN cause 34). chan_sip and PJSIP use the same mapping. It congests the call at once. No timer is involved. The carrier answered; the answer was “no capacity.”

Under chan_sip, both paths land on a congestion-flavored dial status. That is why operators read AUTO-CONGEST as “the carrier sent 503.” But the fixes point in opposite directions:

  • AUTO-CONGEST (no response): a reachability problem. Signaling is blocked by a firewall, a NAT device is eating packets, the proxy address is wrong, or the trunk is dead. The matching symptom is Asterisk’s retransmission-timeout warning. It logs when retransmits of a critical packet run out. Check the network path first, the carrier second.
  • Actual 503 in the SIP trace: a capacity or policy issue at the carrier. Your packets are arriving fine. Work the ladder above.

A SIP trace settles it instantly. A visible 503 on the wire is a capacity problem. INVITEs going out with nothing coming back is a reachability problem, and no pacing change will fix it.

What PJSIP does instead

PJSIP has no auto-congest timer. The pjproject transaction layer runs its own timer B, also 64×T1, or 32 seconds at the default 500 ms T1. When it expires with no response, pjproject ends the INVITE with an internally generated 408 Request Timeout. Asterisk maps that to hangup cause 18 (no user response), not to congestion.

One PJSIP detail catches operators out. pjproject reports a local transport failure, such as a send that fails or a connection that cannot be opened, as an internal 503. A failed DNS lookup becomes an internal 502 instead. So under PJSIP, a 503 in the Asterisk log is not proof the carrier sent one. Run pjsip set logger on and check whether a 503 actually arrived on the wire.

Reading a 503 pattern

One 503 is noise. The pattern is the diagnosis:

  • Is Retry-After present? If the response carries one, the carrier told you exactly how long to back off. Honor it. Hammering a route that just declared itself unavailable makes the event longer.
  • Bursts or constant? Failures that spike when dialing ramps and clear when it slows point at CPS or channel ceilings. A steady failure rate at any pace points at congestion, routing, or account state.
  • Per-destination or global? Pull failed calls by destination prefix. One prefix failing while the rest completes means an upstream route problem. Everything failing at once means capacity, maintenance, or your account.
  • What do the logs say? Confirm which path produced the failure. Was it a 503 that arrived in the SIP trace, a chan_sip auto-congest with no inbound packets at all, or a PJSIP timeout or transport error generated locally? The VICIdial troubleshooting guide walks the wider log-reading workflow.

Fixing it: pacing, spreading, failover, headroom

Four fixes, in the order most teams should apply them:

  1. Pace below your ceiling. Find your trunk’s actual CPS and channel limits. Set the dialer to stay under them. A dialer that bursts over its cap doesn’t complete more calls. It turns the overage into 503s and retries.

  2. Spread across trunks. Splitting campaigns across several trunks keeps a single ceiling from capping the whole floor. It also keeps a per-route failure to a slice of your traffic.

  3. Build failover into the dialplan. 503 is the RFC’s alternate-server signal, so it is the natural trigger for route-advance logic. On 503, try the next trunk rather than burning the attempt. Trunk and dialplan structure for this is covered in the VICIdial carrier setup walkthrough.

  4. Buy actual headroom. Chronic 503s at rates a campaign truly needs are a capacity-shopping signal, not a tuning problem. Before you commit, get the CPS limit your traffic profile will run at in writing from any carrier you are considering, and ask whether concurrent channels are capped. SIPNEX, an FCC-licensed carrier running SIP trunks for dialer operations, does not cap channels at the trunk level, but CPS is a separate limit, so ask us for ours too. Size to your burst rate, not your average.

One last step remains. If every call returns 503 constantly, for days, on every destination, at normal pacing, stop tuning. A pattern like that points at account state, not capacity. The fix is a call to your carrier, not a dialplan change. The suspended dialer account runbook covers what to secure while you have it. If your carrier has gone quiet, talk to us.

Frequently asked questions

Is a SIP 503 the same as Asterisk’s AUTO-CONGEST?

No. They are two different paths with one shared outcome. AUTO-CONGEST is a chan_sip timer that fires when no response of any kind arrives before SIP timer B expires (32 seconds by default). That is a reachability problem. A real 503 is an answer from the carrier, mapped at once to a congestion hangup cause. That is a capacity or policy problem. PJSIP has no auto-congest timer; a silent far end there ends as a locally generated 408. A SIP trace settles which one you have.

Should my dialer retry calls that hit a 503?

Yes, but not right away, and ideally not on the same route. RFC 3261 says the client should try an alternate server; 503 is the standard failover trigger. Honor any Retry-After header. Do not hammer the same trunk at full pace during a capacity event. Every retry above the ceiling is just another 503.

Why do I only get 503 errors during dialing bursts?

You are hitting a rate ceiling, not a route failure. Carriers commonly return 503 when your calls-per-second limit is reached. So failures spike as the dialer ramps and vanish when pacing drops. Pace below your trunk’s actual CPS and channel limits, spread campaigns across trunks, or buy more headroom.


SIPNEX is an FCC-licensed carrier running SIP trunks for dialer operations. Capacity talks happen before you commit, not after the 503s start. See how a trunk is built for burst dialing in the VICIdial SIP trunk guide, or talk to us about your traffic profile.

SIPNEX

The carrier built by operators, for operators.

FCC-licensed carrier with its own STIR/SHAKEN SP certificate. Operator-owned. SIP trunks built for operators who dial at volume.