11.1 NAPTR → SRV → A/AAAA
A SIP URI names a domain, such as sip:bob@biloxi.example, not a server. Before Proxy A can send anything, it needs three things: a transport, a port, and an IP address. RFC 3263 gets them from DNS, with three kinds of records, in this order:
| Record | Question it answers | biloxi.example has |
|---|---|---|
| NAPTR | Which transports does the domain offer, and which does it prefer? | TLS first, then TCP, then UDP |
| SRV | Which servers offer SIP over this transport, on which port? | sip1 and sip2 sharing the load, backup on standby |
| A / AAAA | What is the IP address of each server? | 203.0.113.11, 203.0.113.12, 198.51.100.50 |
biloxi.example. IN NAPTR 50 50 "s" "SIPS+D2T" "" _sips._tcp.biloxi.example.
biloxi.example. IN NAPTR 90 50 "s" "SIP+D2T" "" _sip._tcp.biloxi.example.
biloxi.example. IN NAPTR 100 50 "s" "SIP+D2U" "" _sip._udp.biloxi.example.
_sip._udp.biloxi.example. IN SRV 10 60 5060 sip1.biloxi.example.
_sip._udp.biloxi.example. IN SRV 10 40 5060 sip2.biloxi.example.
_sip._udp.biloxi.example. IN SRV 20 0 5060 backup.biloxi.example.
sip1.biloxi.example. IN A 203.0.113.11
sip2.biloxi.example. IN A 203.0.113.12
backup.biloxi.example. IN A 198.51.100.50
(_sip._tcp and _sips._tcp have the same three servers; _sips._tcp uses port 5061.)
The client resolves the URI that decides the next hop: the first Route value, or the Request-URI when there is no Route (Module 10). DNS gives it somewhere to send the request. It does not change the request:
“The procedures defined here in no way affect this URI (i.e., the URI is not rewritten with the result of the DNS lookup), they only result in an IP address, port and transport protocol where the request can be sent.”Read the section ↗
Step through the lookups below. Change the Request-URI, the transports Proxy A supports, and the zone; then select a server to take it down, and watch the failover. The rest of this module explains each rule.
DNS resolver step-through
Proxy A looks for the servers of biloxi.example
- DNS
- SIP
- Failure
sip:bob@biloxi.exampleNAPTR: SIP+D2T has the lowest order among the services the client supports, so TCP.
biloxi.example. 3600 IN NAPTR 50 50 "s" "SIPS+D2T" "" _sips._tcp.biloxi.example. biloxi.example. 3600 IN NAPTR 90 50 "s" "SIP+D2T" "" _sip._tcp.biloxi.example. biloxi.example. 3600 IN NAPTR 100 50 "s" "SIP+D2U" "" _sip._udp.biloxi.example.
The client keeps the services it supports, and takes the lowest order: SIP+D2T.
11.2 SRV priority and weight: failover and load sharing
Each SRV record has two numbers that decide the order in which the client tries the servers. Priority is for failover: lower first, always.
“A client MUST attempt to contact the target host with the lowest-numbered priority it can reach; target hosts with the same priority SHOULD be tried in an order defined by the weight field.”Read the section ↗
Weight is for load sharing among servers with the same priority:
“The weight field specifies a relative weight for entries with the same priority. Larger weights SHOULD be given a proportionately higher probability of being selected.”Read the section ↗
RFC 2782 turns the weights into an order with a random draw. Put the records of one priority in a list, add up the weights as a running sum, draw a random number from 0 to the total, and take the first record whose running sum reaches it. Then repeat with the records that are left.
| Record | Weight | Running sum | Picked first when the draw is |
|---|---|---|---|
| sip1 | 60 | 60 | 0–60 |
| sip2 | 40 | 100 | 61–100 |
| backup | 0 | — | never first: priority 20 waits until sip1 and sip2 both fail |
Over many requests, sip1 gets about 60% and sip2 about 40%. Select Draw again in the step-through above to see a new draw. Weight 0 means “no preference”:
“Domain administrators SHOULD use Weight 0 when there isn't any server selection to do, to make the RR easier to read for humans (less noisy). In the presence of records containing weights greater than 0, records with weight 0 should have a very small chance of being selected.”Read the section ↗
The numbers in DNS are static. They say “this server is bigger”, not “this server is busy now”: DNS answers are cached for the TTL, and the client cannot know the current load.
“Weight is only intended for static, not dynamic, server selection.”Read the section ↗
| You want | Use |
|---|---|
| Active and standby | Different priorities: 10 for the active server, 20 for the standby |
| Two equal servers | One priority, equal weights (50 and 50) |
| One server twice as big | One priority, weights 2 and 1 |
| No SIP over this transport | One SRV record with the target . |
“A Target of "." means that the service is decidedly not available at this domain.”Read the section ↗
11.3 How the client selects a transport
The client chooses the transport first, with the first rule that applies:
| The URI has… | Transport | Then for the address |
|---|---|---|
;transport= |
The one it names | SRV for that transport |
| An IP address as host | UDP (TLS for a SIPS URI) | No DNS at all |
A port, such as :5060 |
UDP (TLS for a SIPS URI) | A/AAAA of the domain, at that port |
| A domain only | The most preferred NAPTR service that the client supports | SRV named in the NAPTR record |
| A domain only, and no NAPTR | The first transport with SRV records | Those SRV records |
| A domain only, no NAPTR, no SRV | UDP (TLS for a SIPS URI) | A/AAAA of the domain, default port |
“If the URI specifies a transport protocol in the transport parameter, that transport protocol SHOULD be used.”Read the section ↗
“Otherwise, if no transport protocol is specified, but the TARGET is a numeric IP address, the client SHOULD use UDP for a SIP URI, and TCP for a SIPS URI. Similarly, if no transport protocol is specified, and the TARGET is not numeric, but an explicit port is provided, the client SHOULD use UDP for a SIP URI, and TCP for a SIPS URI.”Read the section ↗
“Otherwise, if no transport protocol or port is specified, and the target is not a numeric IP address, the client SHOULD perform a NAPTR query for the domain in the URI.”Read the section ↗
The NAPTR service names the transport: SIP+D2U is UDP, SIP+D2T is TCP, and SIPS+D2T is TLS over TCP. The client drops the services it cannot use and takes the lowest order. In the example, a client that supports UDP and TCP picks TCP, because SIP+D2T has order 90 and SIP+D2U 100. A client with TLS takes SIPS+D2T, even for a SIP URI:
“First, a client resolving a SIPS URI MUST discard any services that do not contain "SIPS" as the protocol in the service field. The converse is not true, however. A client resolving a SIP URI SHOULD retain records with "SIPS" as the protocol, if the client supports TLS.”Read the section ↗
A SIP server that publishes NAPTR must list at least the three services:
“If a SIP proxy, redirect server, or registrar is to be contacted through the lookup of NAPTR records, there MUST be at least three records - one with a "SIP+D2T" service field, one with a "SIP+D2U" service field, and one with a "SIPS+D2T" service field.”Read the section ↗
Many domains publish no NAPTR records, only SRV. The client then asks for each transport it supports, and without SRV, it uses the A record:
“If no NAPTR records are found, the client constructs SRV queries for those transport protocols it supports, and does a query for each.”Read the section ↗
“If no SRV records are found, the client SHOULD use TCP for a SIPS URI, and UDP for a SIP URI.”Read the section ↗
“If no SRV records were found, the client performs an A or AAAA record lookup of the domain name. The result will be a list of IP addresses, each of which can be contacted using the transport protocol determined previously, at the default port for that transport.”Read the section ↗
The transport from DNS is not the last word. A large request — over 1300 bytes when the path MTU is unknown — must use a congestion-controlled transport such as TCP, even when DNS chose UDP (Modules 2 and 19):
“If a request is within 200 bytes of the path MTU, or if it is larger than 1300 bytes and the path MTU is unknown, the request MUST be sent using an RFC 2914 [43] congestion controlled transport protocol, such as TCP.”Read the section ↗
11.4 Failover on 503 and on transport errors
The client tries the first address in its list. It moves to the next one when the attempt fails:
“For SIP requests, failure occurs if the transaction layer reports a 503 error response or a transport failure of some sort (generally, due to fatal ICMP errors in UDP or connection failures in TCP). Failure also occurs if the transaction layer times out without ever having received any response, provisional or final (i.e., timer B or timer F in RFC 3261 [1] fires).”Read the section ↗
| What happens | The client moves on… |
|---|---|
| The server answers 503 Service Unavailable | At once |
| An ICMP port unreachable (UDP), or the TCP connection is refused | At once |
| Nothing answers at all | When Timer B (INVITE) or Timer F (other requests) fires: 32 seconds |
| Any other response — 100 Trying, a 4xx, a 2xx | Never: the server has the request |
Each new try is a new transaction, with a new branch:
“If a failure occurs, the client SHOULD create a new request, which is identical to the previous, but has a different value of the Via branch ID than the previous (and therefore constitutes a new SIP transaction). That request is sent to the next element in the list as specified by RFC 2782.”Read the section ↗
Once a server answers, the transaction stays there. Retransmissions, the ACK for a non-2xx, and CANCEL go to the same address. The ACK for a 2xx is a new transaction, so it follows the route set instead (Module 10):
“The procedures here MUST be done exactly once per transaction, where transaction is as defined in [1]. That is, once a SIP server has successfully been contacted (success is defined below), all retransmissions of the SIP request and the ACK for non-2xx SIP responses to INVITE MUST be sent to the same host. Furthermore, a CANCEL for a particular SIP request MUST be sent to the same SIP server that the SIP request was delivered to.”Read the section ↗
“Because the ACK request for 2xx responses to INVITE constitutes a different transaction, there is no requirement that it be delivered to the same server that received the original request (indeed, if that server did not record-route, it will not get the ACK).”Read the section ↗
A server that is switched off is the slow case: the caller waits 32 seconds in silence before the second server rings. That is why many proxies use a shorter timer for each DNS target, and remember a failed address for a while, so that the next call goes straight to a server that works. A server that is up but overloaded should answer 503, so that clients move on at once.
Call flow · NAPTR, SRV, and failover
Fixed: SRV lists a second server
- SIP
- DNS / STUN / ICE
- ⚠Problem
Alice calls Bob. Her phone sends the INVITE to Proxy A, her outbound proxy.
All steps as text
- Alice → Proxy A: INVITE. Alice calls Bob. Her phone sends the INVITE to Proxy A, her outbound proxy.
- Proxy A → Alice: 100 Trying. Proxy A answers 100 Trying on this hop, before it looks up biloxi.example.
- Proxy A → DNS resolver: DNS query (NAPTR). The Request-URI has a domain, no port, and no transport. Proxy A asks which transports biloxi.example offers.
- DNS resolver → Proxy A: DNS answer (NAPTR). Three services: TLS, TCP, and UDP, in that order. Proxy A supports only UDP, so it keeps SIP+D2U.
- Proxy A → DNS resolver: DNS query (SRV). Proxy A asks for the UDP servers: the SRV name from the NAPTR record.
- DNS resolver → Proxy A: DNS answer (SRV). Two servers at priority 10 share the load 60:40; a backup waits at priority 20. The addresses come in the same answer.
- Proxy A → Proxy B1: INVITE. The weight draw picks sip1 first. Proxy A sends the INVITE to 203.0.113.11, but sip1 is down.
- Proxy A → Proxy B1: INVITE (again). Proxy A retransmits at intervals of 0.5, 1, 2, 4, 8, and 16 seconds. Nothing answers.
- Proxy A → Proxy B2: INVITE. After 32 seconds Timer B fires. Proxy A sends a new INVITE, with a new branch, to the next server: sip2. Problem: Alice heard nothing for 32 seconds. Many proxies use a shorter timer for this reason.
- Proxy B2 → Proxy A: 100 Trying. sip2 answers. Retransmissions and a CANCEL for this INVITE now go only to sip2.
- Proxy B2 → Bob: INVITE. sip2 finds Bob's Contact and forwards the INVITE.
- Bob → Proxy B2: 180 Ringing. Bob's phone rings.
- Proxy B2 → Proxy A: 180 Ringing. sip2 forwards the 180 Ringing to Proxy A.
- Proxy A → Alice: 180 Ringing. Alice hears ringing, 32 seconds late. The call works only because DNS listed a second server.
Common mistakes
A port in the URI stops the SRV lookup
A trunk or route configured as sip:biloxi.example:5060 looks harmless: 5060 is the SIP port anyway. But a port in the URI tells the client to skip NAPTR and SRV and to look up only the A record of the domain. All the priorities, weights, and backup servers in SRV are never seen. When the one address in the A record is down, the call fails — while sip2 is up.
“This is a change from RFC 2543. Previously, if the port was explicit, but with a value of 5060, SRV records were used. Now, A or AAAA records will be used.”Read the section ↗
Broken
Broken: a port in the URI stops SRV
- SIP
- DNS / STUN / ICE
- ⚠Problem
Alice calls Bob. Her phone sends the INVITE to Proxy A, her outbound proxy.
All steps as text
- Alice → Proxy A: INVITE. Alice calls Bob. Her phone sends the INVITE to Proxy A, her outbound proxy.
- Proxy A → Alice: 100 Trying. Proxy A answers 100 Trying on this hop, before it looks up biloxi.example.
- Proxy A → DNS resolver: DNS query (A). Proxy A's route to biloxi.example has a port: sip:biloxi.example:5060. With a port, it asks only for A records.
- DNS resolver → Proxy A: DNS answer (A). One address: 203.0.113.11, which is sip1. Proxy A never sees sip2 or the backup in SRV.
- Proxy A → Proxy B1: INVITE. Proxy A sends the INVITE to 203.0.113.11, port 5060. sip1 is down.
- Proxy A → Proxy B1: INVITE (again). Proxy A retransmits for 32 seconds. Nothing answers.
- Proxy A → Alice: 408 Request Timeout. Timer B fires. Proxy A has no other address to try, so the call fails. Problem: sip2 was up the whole time. The port in the URI hid it.
- Alice → Proxy A: ACK. Alice's phone ACKs the 408 on this hop.
Remember: configure the domain with no port, as sip:biloxi.example. Add a port or a transport only when the other side asks for it — and then check that its SRV records are not needed.
Hard-coding IP addresses
sip:bob@203.0.113.11 needs no DNS at all: the client sends to that address, with UDP, on port 5060. It works until that server is down, replaced, or moved — and then there is nothing to fail over to. A carrier that adds a server, changes an address, or moves traffic during maintenance does it in DNS; a hard-coded address sees none of it. TLS also needs the name: the certificate names the domain, not the address (Module 14).
Remember: if a trace shows every request going to one address, and the calls fail when that server is down, look at the configuration for a hard-coded address or a port.
An SRV target that is an alias, or has no address
The SRV target must be a host name with its own A or AAAA records. A target that is a CNAME — or a name with no address records — breaks the rule below. Some clients follow the alias anyway and others give up, so the failure can appear on some clients only.
“There MUST be one or more address records for this name, the name MUST NOT be an alias (in the sense of RFC 1034 or RFC 2181).”Read the section ↗
Remember: check every SRV target with dig A and dig AAAA, and make sure it returns addresses directly, not a CNAME.
Summary
- A client resolves the first Route value, or the Request-URI. DNS gives an address, a port, and a transport; it never changes the URI.
- NAPTR chooses the transport; SRV lists the servers and ports; A/AAAA give the addresses.
- Priority is for failover: the lowest number first. Weight shares the load among servers of the same priority, with a random draw.
- The transport comes from
;transport=, else from the defaults for an IP address or a port, else from NAPTR, else from SRV, else UDP (TLS for SIPS). - The client moves to the next server after a 503, a transport error, or no answer at all — the last takes 32 seconds. Each try is a new transaction with a new branch.
- A port or an IP address in the URI turns off SRV, and with it all failover.