STRESS_README.html revision 1.1.1.6 1 <!doctype html public "-//W3C//DTD HTML 4.01 Transitional//EN"
2 "http://www.w3.org/TR/html4/loose.dtd">
3
4 <html>
5
6 <head>
7
8 <title>Postfix Stress-Dependent Configuration</title>
9
10 <meta http-equiv="Content-Type" content="text/html; charset=utf-8">
11
12 </head>
13
14 <body>
15
16 <h1><img src="postfix-logo.jpg" width="203" height="98" ALT="">Postfix
17 Stress-Dependent Configuration</h1>
18
19 <hr>
20
21 <h2>Overview </h2>
22
23 <p> This document describes the symptoms of Postfix SMTP server
24 overload. It presents permanent main.cf changes to avoid overload
25 during normal operation, and temporary main.cf changes to cope with
26 an unexpected burst of mail. This document makes specific suggestions
27 for Postfix 2.5 and later which support stress-adaptive behavior,
28 and for earlier Postfix versions that don't. </p>
29
30 <p> Topics covered in this document: </p>
31
32 <ul>
33
34 <li><a href="#overload"> Symptoms of Postfix SMTP server overload </a>
35
36 <li><a href="#adapt"> Automatic stress-adaptive behavior </a>
37
38 <li><a href="#concurrency"> Service more SMTP clients at the same time </a>
39
40 <li><a href="#time"> Spend less time per SMTP client </a>
41
42 <li><a href="#hangup"> Disconnect suspicious SMTP clients </a>
43
44 <li><a href="#legacy"> Temporary measures for older Postfix releases </a>
45
46 <li><a href="#feature"> Detecting support for stress-adaptive behavior </a>
47
48 <li><a href="#forcing"> Forcing stress-adaptive behavior on or off </a>
49
50 <li><a href="#other"> Other measures to off-load zombies </a>
51
52 <li><a href="#credits"> Credits </a>
53
54 </ul>
55
56 <h2><a name="overload"> Symptoms of Postfix SMTP server overload </a></h2>
57
58 <p> Under normal conditions, the Postfix SMTP server responds
59 immediately when an SMTP client connects to it; the time to deliver
60 mail is noticeable only with large messages. Performance degrades
61 dramatically when the number of SMTP clients exceeds the number of
62 Postfix SMTP server processes. When an SMTP client connects while
63 all Postfix SMTP server processes are busy, the client must wait
64 until a server process becomes available. </p>
65
66 <p> SMTP server overload may be caused by a surge of legitimate
67 mail (example: a DNS registrar opens a new zone for registrations),
68 by mistake (mail explosion caused by a forwarding loop) or by malice
69 (worm outbreak, botnet, or other illegitimate activity). </p>
70
71 <p> Symptoms of Postfix SMTP server overload are: </p>
72
73 <ul>
74
75 <li> <p> Remote SMTP clients experience a long delay before Postfix
76 sends the "220 hostname.example.com ESMTP Postfix" greeting. </p>
77
78 <ul>
79
80 <li> <p> NOTE: Broken DNS configurations can also cause lengthy
81 delays before Postfix sends "220 hostname.example.com ...". These
82 delays also exist when Postfix is NOT overloaded. </p>
83
84 <li> <p> NOTE: To avoid "overload" delays for end-user mail
85 clients, enable the "submission" service entry in master.cf (present
86 since Postfix 2.1), and tell users to connect to this instead of
87 the public SMTP service. </p>
88
89 </ul>
90
91 <li> <p> The Postfix SMTP server logs an increased number of "lost
92 connection after CONNECT" events. This happens because remote SMTP
93 clients disconnect before Postfix answers the connection. </p>
94
95 <ul>
96
97 <li> <p> NOTE: A portscan for open SMTP ports can also result in
98 "lost connection ..." logfile messages. </p>
99
100 </ul>
101
102 <li> <p> Postfix 2.3 and later logs a warning that all server ports
103 are busy: </p>
104
105 <pre>
106 Oct 3 20:39:27 spike postfix/master[28905]: warning: service "smtp"
107 (25) has reached its process limit "30": new clients may experience
108 noticeable delays
109 Oct 3 20:39:27 spike postfix/master[28905]: warning: to avoid this
110 condition, increase the process count in master.cf or reduce the
111 service time per client
112 Oct 3 20:39:27 spike postfix/master[28905]: warning: see
113 <a href="http://www.postfix.org/STRESS_README.html">http://www.postfix.org/STRESS_README.html</a> for examples of
114 stress-adapting configuration settings
115 </pre>
116
117 </ul>
118
119 <p> Legitimate mail that doesn't get through during an episode of
120 Postfix SMTP server overload is not necessarily lost. It should
121 still arrive once the situation returns to normal, as long as the
122 overload condition is temporary. </p>
123
124 <h2><a name="adapt"> Automatic stress-adaptive behavior </a></h2>
125
126 <p> Postfix version 2.5 introduces automatic stress-adaptive behavior.
127 It works as follows. When a "public" network service such as the
128 SMTP server runs into an "all server ports are busy" condition, the
129 Postfix master(8) daemon logs a warning, restarts the service
130 (without interrupting existing network sessions), and runs the
131 service with "-o stress=yes" on the server process command line:
132 </p>
133
134 <blockquote>
135 <pre>
136 80821 ?? S 0:00.24 smtpd -n smtp -t inet -u -c -o stress=yes
137 </pre>
138 </blockquote>
139
140 <p> Normally, the Postfix master(8) daemon runs such a service with
141 "-o stress=" on the command line (i.e. with an empty parameter
142 value): </p>
143
144 <blockquote>
145 <pre>
146 83326 ?? S 0:00.28 smtpd -n smtp -t inet -u -c -o stress=
147 </pre>
148 </blockquote>
149
150 <p> You won't see "-o stress" command-line parameters with services
151 that have local clients only. These include services internal to
152 Postfix such as the queue manager, and services that listen on a
153 loopback interface only, such as after-filter SMTP services. </p>
154
155 <p> The "stress" parameter value is the key to making main.cf
156 parameter settings stress adaptive. The following settings are the
157 default with Postfix 2.6 and later. </p>
158
159 <blockquote>
160 <pre>
161 1 smtpd_timeout = ${stress?{10}:{300}}s
162 2 smtpd_hard_error_limit = ${stress?{1}:{20}}
163 3 smtpd_junk_command_limit = ${stress?{1}:{100}}
164 4 # Parameters added after Postfix 2.6:
165 5 smtpd_per_record_deadline = ${stress?{yes}:{no}}
166 6 smtpd_starttls_timeout = ${stress?{10}:{300}}s
167 7 address_verify_poll_count = ${stress?{1}:{3}}
168 </pre>
169 </blockquote>
170
171 <p> Postfix versions before 3.0 use the older form ${stress?x}${stress:y}
172 instead of the newer form ${stress?{x}:{y}}. </p>
173
174 <p> The syntax of ${name?{value}:{value}}, ${name?value} and
175 ${name:value} is explained at the beginning of the postconf(5)
176 manual page. </p>
177
178 <p> Translation: <p>
179
180 <ul>
181
182 <li> <p> Line 1: under conditions of stress, use an smtpd_timeout
183 value of 10 seconds instead of the default 300 seconds. Experience
184 on the postfix-users list from a variety of sysadmins shows that
185 reducing the "normal" smtpd_timeout to 60s is unlikely to affect
186 legitimate clients. However, it is unlikely to become the Postfix
187 default because it's not RFC compliant. Setting smtpd_timeout to
188 10s or even 5s under stress will still allow most
189 legitimate clients to connect and send mail, but may delay mail
190 from some clients. No mail should be lost, as long as this measure
191 is used only temporarily. </p>
192
193 <li> <p> Line 2: under conditions of stress, use an smtpd_hard_error_limit
194 of 1 instead of the default 20. This disconnects clients
195 after a single error, giving other clients a chance to connect.
196 However, this may cause significant delays with legitimate mail,
197 such as a mailing list that contains a few no-longer-active user
198 names that didn't bother to unsubscribe. No mail should be lost,
199 as long as this measure is used only temporarily. </p>
200
201 <li> <p> Line 3: under conditions of stress, use an
202 smtpd_junk_command_limit of 1 instead of the default 100. This
203 prevents clients from keeping connections open by repeatedly
204 sending HELO, EHLO, NOOP, RSET, VRFY or ETRN commands. </p>
205
206 <li> <p> Line 5: under conditions of stress, change the behavior
207 of smtpd_timeout and smtpd_starttls_timeout, from a time limit per
208 read or write system call, to a time limit to send or receive a
209 complete record (an SMTP command line, SMTP response line, SMTP
210 message content line, or TLS protocol message). </p>
211
212 <li> <p> Line 6: under conditions of stress, reduce the time limit
213 for TLS protocol handshake messages to 10 seconds, from the default
214 value of 300 seconds. See also the smtpd_timeout discussion above.
215 </p>
216
217 <li> <p> Line 7: under conditions of stress, do not wait up to 6
218 seconds for the completion of an address verification probe. If the
219 result is not already in the address verification cache, reply
220 immediately with $unverified_recipient_tempfail_action or
221 $unverified_sender_tempfail_action. No mail should be lost, as long
222 as this measure is used only temporarily. </p>
223
224 </ul>
225
226 <p> NOTE: Please keep in mind that the stress-adaptive feature is
227 a fairly desperate measure to keep <b>some</b> legitimate mail
228 flowing under overload conditions. If a site is reaching the SMTP
229 server process limit when there isn't an attack or bot flood
230 occurring, then either the process limit needs to be raised or more
231 hardware needs to be added. </p>
232
233 <h2><a name="concurrency"> Service more SMTP clients at the same time </a> </h2>
234
235 <p> This section and the ones that follow discuss permanent measures
236 against mail server overload. </p>
237
238 <p> One measure to avoid the "all server processes busy" condition
239 is to service more SMTP clients simultaneously. For this you need
240 to increase the number of Postfix SMTP server processes. This will
241 improve the
242 responsiveness for remote SMTP clients, as long as the server machine
243 has enough hardware and software resources to run the additional
244 processes, and as long as the file system can keep up with the
245 additional load. </p>
246
247 <ul>
248
249 <li> <p> You increase the number of SMTP server processes either
250 by increasing the default_process_limit in main.cf (line 3 below),
251 or by increasing the SMTP server's "maxproc" field in master.cf
252 (line 10 below). Either way, you need to issue a "postfix reload"
253 command to make the change effective. </p>
254
255 <li> <p> Process limits above 1000 require Postfix version 2.4 or
256 later, and an operating system that supports kernel-based event
257 filters (BSD kqueue(2), Linux epoll(4), or Solaris /dev/poll).
258 </p>
259
260 <li> <p> More processes use more memory. You can reduce the Postfix
261 memory footprint by using cdb:
262 lookup tables instead of Berkeley DB's hash: or btree: tables. </p>
263
264 <pre>
265 1 /etc/postfix/main.cf:
266 2 # Raise the global process limit, 100 since Postfix 2.0.
267 3 default_process_limit = 200
268 4
269 5 /etc/postfix/master.cf:
270 6 # =============================================================
271 7 # service type private unpriv chroot wakeup maxproc command
272 8 # =============================================================
273 9 # Raise the SMTP service process limit only.
274 10 smtp inet n - n - 200 smtpd
275 </pre>
276
277 <li> <p> NOTE: older versions of the SMTPD_POLICY_README document
278 contain a mistake: they configure a fixed number of policy daemon
279 processes. When you raise the SMTP server's "maxproc" field in
280 master.cf, SMTP server processes will report problems when connecting
281 to policy server processes, because there aren't enough of them.
282 Examples of errors are "connection refused" or "operation timed
283 out". </p>
284
285 <p> To fix, edit master.cf and specify a zero "maxproc" field
286 in all policy server entries; see line 6 in the example below.
287 Issue a "postfix reload" command to make the change effective. </p>
288
289 <pre>
290 1 /etc/postfix/master.cf:
291 2 # =============================================================
292 3 # service type private unpriv chroot wakeup maxproc command
293 4 # =============================================================
294 5 # Disable the policy service process limit.
295 6 policy unix - n n - 0 spawn
296 7 user=nobody argv=/some/where/policy-server
297 </pre>
298
299 </ul>
300
301 <h2><a name="time"> Spend less time per SMTP client </a></h2>
302
303 <p> When increasing the number of SMTP server processes is not
304 practical, you can improve Postfix server responsiveness by eliminating
305 delays. When Postfix spends less time per SMTP session, the same
306 number of SMTP server processes can service more clients in a given
307 amount of time. </p>
308
309 <ul>
310
311 <li> <p> Eliminate non-functional RBL lookups (blocklists that are
312 no longer in operation). These lookups can degrade performance.
313 Postfix logs a warning when an RBL server does not respond. </p>
314
315 <li> <p> Eliminate redundant RBL lookups (people often use multiple
316 Spamhaus RBLs that include each other). To find out whether RBLs
317 include other RBLs, look up the websites that document the RBL's
318 policies. </p>
319
320 <li> <p> Eliminate header_checks and body_checks, and keep just a few
321 emergency patterns to block the latest worm explosion or backscatter
322 mail. See BACKSCATTER_README for examples of the latter.
323
324 <li> <p> Group your header_checks and body_checks patterns to avoid
325 unnecessary pattern matching operations:
326
327 <pre>
328 1 /etc/postfix/header_checks:
329 2 if /^Subject:/
330 3 /^Subject: virus found in mail from you/ reject
331 4 /^Subject: ..other../ reject
332 5 endif
333 6
334 7 if /^Received:/
335 8 /^Received: from (postfix\.org) / reject forged client name in received header: $1
336 9 /^Received: from ..other../ reject ....
337 10 endif
338 </pre>
339
340 </ul>
341
342 <h2><a name="hangup"> Disconnect suspicious SMTP clients </a></h2>
343
344 <p> Under conditions of overload you can improve Postfix SMTP server
345 responsiveness by hanging up on suspicious clients, so that other
346 clients get a chance to talk to Postfix. </p>
347
348 <ul>
349
350 <li> <p> Use "521" SMTP reply codes (Postfix 2.6 and later) or "421"
351 (Postfix 2.3-2.5) to hang up on clients that that match botnet-related
352 RBLs (see next bullet) or that match selected non-RBL restrictions
353 such as SMTP access maps. The Postfix SMTP server will reject mail
354 and disconnect without waiting for the remote SMTP client to send
355 a QUIT command. </p>
356
357 <li> <p> To hang up connections from denylisted zombies, you can
358 set specific Postfix SMTP server reject codes for specific RBLs,
359 and for individual responses from specific RBLs. We'll use
360 zen.spamhaus.org as an example; by the time you read this document,
361 details may have changed. Right now, their documents say that a
362 response of 127.0.0.10 or 127.0.0.11 indicates a dynamic client IP
363 address, which means that the machine is probably running a bot of
364 some kind. To give a 521 response instead of the default 554
365 response, use something like: </p>
366
367 <pre>
368 1 /etc/postfix/main.cf:
369 2 smtpd_client_restrictions =
370 3 permit_mynetworks
371 4 reject_rbl_client zen.spamhaus.org=127.0.0.10
372 5 reject_rbl_client zen.spamhaus.org=127.0.0.11
373 6 reject_rbl_client zen.spamhaus.org
374 7
375 8 rbl_reply_maps = hash:/etc/postfix/rbl_reply_maps
376 9
377 10 /etc/postfix/rbl_reply_maps:
378 11 # With Postfix 2.3-2.5 use "421" to hang up connections.
379 12 zen.spamhaus.org=127.0.0.10 521 4.7.1 Service unavailable;
380 13 $rbl_class [$rbl_what] blocked using
381 14 $rbl_domain${rbl_reason?; $rbl_reason}
382 15
383 16 zen.spamhaus.org=127.0.0.11 521 4.7.1 Service unavailable;
384 17 $rbl_class [$rbl_what] blocked using
385 18 $rbl_domain${rbl_reason?; $rbl_reason}
386 </pre>
387
388 <p> Although the above example shows three RBL lookups (lines 4-6),
389 Postfix will only do a single DNS query, so it does not affect the
390 performance. </p>
391
392 <li> <p> With Postfix 2.3-2.5, use reply code 421 (521 will not
393 cause Postfix to disconnect). The down-side of replying with 421
394 is that it works only for zombies and other malware. If the client
395 is running a real MTA, then it may connect again several times until
396 the mail expires in its queue. When this is a problem, stick with
397 the default 554 reply, and use "smtpd_hard_error_limit = 1" as
398 described below. </p>
399
400 <li> <p> You can automatically turn on the above overload measure
401 with Postfix 2.5 and later, or with earlier releases that contain
402 the stress-adaptive behavior source code patch from the mirrors
403 listed at http://www.postfix.org/download.html. Simply replace line
404 above 8 with: </p>
405
406 <pre>
407 8 rbl_reply_maps = ${stress?hash:/etc/postfix/rbl_reply_maps}
408 </pre>
409
410 </ul>
411
412 <p> More information about automatic stress-adaptive behavior is
413 in section "<a href="#adapt">Automatic stress-adaptive behavior</a>".
414 </p>
415
416 <h2><a name="legacy"> Temporary measures for older Postfix releases </a></h2>
417
418 <p> See the section "<a href="#adapt">Automatic stress-adaptive
419 behavior</a>" if you are running Postfix version 2.5 or later, or
420 if you have applied the source code patch for stress-adaptive
421 behavior from the mirrors listed at http://www.postfix.org/download.html.
422 </p>
423
424 <p> The following measures can be applied temporarily during overload.
425 They still allow <b>most</b> legitimate clients to connect and send
426 mail, but may affect some legitimate clients. </p>
427
428 <ul>
429
430 <li> <p> Reduce smtpd_timeout (default: 300s). Experience on the
431 postfix-users list from a variety of sysadmins shows that reducing
432 the "normal" smtpd_timeout to 60s is unlikely to affect legitimate
433 clients. However, it is unlikely to become the Postfix default
434 because it's not RFC compliant. Setting smtpd_timeout to 10s (line
435 2 below) or even 5s under stress will still allow <b>most</b>
436 legitimate clients to connect and send mail, but may delay mail
437 from some clients. No mail should be lost, as long as this measure
438 is used only temporarily. </p>
439
440 <li> <p> Reduce smtpd_hard_error_limit (default: 20). Setting this
441 to 1 under stress (line 3 below) helps by disconnecting clients
442 after a single error, giving other clients a chance to connect.
443 However, this may cause significant delays with legitimate mail,
444 such as a mailing list that contains a few no-longer-active user
445 names that didn't bother to unsubscribe. No mail should be lost,
446 as long as this measure is used only temporarily. </p>
447
448 <li> <p> Use an smtpd_junk_command_limit of 1 instead of the default
449 100. This prevents clients from keeping idle connections open by
450 repeatedly sending NOOP or RSET commands. </p>
451
452 </ul>
453
454 <blockquote>
455 <pre>
456 1 /etc/postfix/main.cf:
457 2 smtpd_timeout = 10
458 3 smtpd_hard_error_limit = 1
459 4 smtpd_junk_command_limit = 1
460 </pre>
461 </blockquote>
462
463 <p> With these measures, no mail should be lost, as long
464 as these measures are used only temporarily. The next section of
465 this document introduces a way to automate this process. </p>
466
467 <h2><a name="feature"> Detecting support for stress-adaptive behavior </a></h2>
468
469 <p> To find out if your Postfix installation supports stress-adaptive
470 behavior, use the "ps" command, and look for the smtpd processes.
471 Postfix has stress-adaptive support when you see "-o stress=" or
472 "-o stress=yes" command-line options. Remember that Postfix never
473 enables stress-adaptive behavior on servers that listen on local
474 addresses only. </p>
475
476 <p> The following example is for FreeBSD or Linux. On Solaris, HP-UX
477 and other System-V flavors, use "ps -ef" instead of "ps ax". </p>
478
479 <blockquote>
480 <pre>
481 $ ps ax|grep smtpd
482 83326 ?? S 0:00.28 smtpd -n smtp -t inet -u -c -o stress=
483 84345 ?? Ss 0:00.11 /usr/bin/perl /usr/libexec/postfix/smtpd-policy.pl
484 </pre>
485 </blockquote>
486
487 <p> You can't use postconf(1) to detect stress-adaptive support.
488 The postconf(1) command ignores the existence of the stress parameter
489 in main.cf, because the parameter has no effect there. Command-line
490 "-o parameter" settings always take precedence over main.cf parameter
491 settings. <p>
492
493 <p> If you configure stress-adaptive behavior in main.cf when it
494 isn't supported, nothing bad will happen. The processes will run
495 as if the stress parameter always has an empty value. </p>
496
497 <h2><a name="forcing"> Forcing stress-adaptive behavior on or off </a></h2>
498
499 <p> You can manually force stress-adaptive behavior on, by adding
500 a "-o stress=yes" command-line option in master.cf. This can be
501 useful for testing overrides on the SMTP service. Issue "postfix
502 reload" to make the change effective. </p>
503
504 <p> Note: setting the stress parameter in main.cf has no effect for
505 services that accept remote connections. </p>
506
507 <blockquote>
508 <pre>
509 1 /etc/postfix/master.cf:
510 2 # =============================================================
511 3 # service type private unpriv chroot wakeup maxproc command
512 4 # =============================================================
513 5 #
514 6 smtp inet n - n - - smtpd
515 7 -o stress=yes
516 8 -o . . .
517 </pre>
518 </blockquote>
519
520 <p> To permanently force stress-adaptive behavior off with a specific
521 service, specify "-o stress=" on its master.cf command line. This
522 may be desirable for the "submission" service. Issue "postfix reload"
523 to make the change effective. </p>
524
525 <p> Note: setting the stress parameter in main.cf has no effect for
526 services that accept remote connections. </p>
527
528 <blockquote>
529 <pre>
530 1 /etc/postfix/master.cf:
531 2 # =============================================================
532 3 # service type private unpriv chroot wakeup maxproc command
533 4 # =============================================================
534 5 #
535 6 submission inet n - n - - smtpd
536 7 -o stress=
537 8 -o . . .
538 </pre>
539 </blockquote>
540
541 <h2><a name="other"> Other measures to off-load zombies </a> </h2>
542
543 <p> The postscreen(8) daemon, introduced with Postfix 2.8, provides
544 additional protection against mail server overload. One postscreen(8)
545 process handles multiple inbound SMTP connections, and decides which
546 clients may talk to a Postfix SMTP server process. By keeping
547 spambots away, postscreen(8) leaves more SMTP server processes
548 available for legitimate clients, and delays the onset of server
549 overload conditions. </p>
550
551 <h2><a name="credits"> Credits </a></h2>
552
553 <ul>
554
555 <li> Thanks to the postfix-users mailing list members for sharing
556 early experiences with the stress-adaptive feature.
557
558 <li> The RBL example and several other paragraphs of text were
559 adapted from postfix-users postings by Noel Jones.
560
561 <li> Wietse implemented stress-adaptive behavior as the smallest
562 possible patch while he should be working on other things.
563
564 </ul>
565
566 </body> </html>
567