README revision 1.1 1 1.1 mrg Copyright 2000, 2001 Free Software Foundation, Inc.
2 1.1 mrg
3 1.1 mrg This file is part of the GNU MP Library.
4 1.1 mrg
5 1.1 mrg The GNU MP Library is free software; you can redistribute it and/or modify
6 1.1 mrg it under the terms of the GNU Lesser General Public License as published by
7 1.1 mrg the Free Software Foundation; either version 3 of the License, or (at your
8 1.1 mrg option) any later version.
9 1.1 mrg
10 1.1 mrg The GNU MP Library is distributed in the hope that it will be useful, but
11 1.1 mrg WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY
12 1.1 mrg or FITNESS FOR A PARTICULAR PURPOSE. See the GNU Lesser General Public
13 1.1 mrg License for more details.
14 1.1 mrg
15 1.1 mrg You should have received a copy of the GNU Lesser General Public License
16 1.1 mrg along with the GNU MP Library. If not, see http://www.gnu.org/licenses/.
17 1.1 mrg
18 1.1 mrg
19 1.1 mrg
20 1.1 mrg
21 1.1 mrg AMD K7 MPN SUBROUTINES
22 1.1 mrg
23 1.1 mrg
24 1.1 mrg This directory contains code optimized for the AMD Athlon CPU.
25 1.1 mrg
26 1.1 mrg The mmx subdirectory has routines using MMX instructions. All Athlons have
27 1.1 mrg MMX, the separate directory is just so that configure can omit it if the
28 1.1 mrg assembler doesn't support MMX.
29 1.1 mrg
30 1.1 mrg
31 1.1 mrg
32 1.1 mrg STATUS
33 1.1 mrg
34 1.1 mrg Times for the loops, with all code and data in L1 cache.
35 1.1 mrg
36 1.1 mrg cycles/limb
37 1.1 mrg mpn_add/sub_n 1.6
38 1.1 mrg
39 1.1 mrg mpn_copyi 0.75 or 1.0 \ varying with data alignment
40 1.1 mrg mpn_copyd 0.75 or 1.0 /
41 1.1 mrg
42 1.1 mrg mpn_divrem_1 17.0 integer part, 15.0 fractional part
43 1.1 mrg mpn_mod_1 17.0
44 1.1 mrg mpn_divexact_by3 8.0
45 1.1 mrg
46 1.1 mrg mpn_l/rshift 1.2
47 1.1 mrg
48 1.1 mrg mpn_mul_1 3.4
49 1.1 mrg mpn_addmul/submul_1 3.9
50 1.1 mrg
51 1.1 mrg mpn_mul_basecase 4.42 cycles/crossproduct (approx)
52 1.1 mrg mpn_sqr_basecase 2.3 cycles/crossproduct (approx)
53 1.1 mrg or 4.55 cycles/triangleproduct (approx)
54 1.1 mrg
55 1.1 mrg Prefetching of sources hasn't yet been tried.
56 1.1 mrg
57 1.1 mrg
58 1.1 mrg
59 1.1 mrg NOTES
60 1.1 mrg
61 1.1 mrg cmov, MMX, 3DNow and some extensions to MMX and 3DNow are available.
62 1.1 mrg
63 1.1 mrg Write-allocate L1 data cache means prefetching of destinations is unnecessary.
64 1.1 mrg
65 1.1 mrg Floating point multiplications can be done in parallel with integer
66 1.1 mrg multiplications, but there doesn't seem to be any way to make use of this.
67 1.1 mrg
68 1.1 mrg Unsigned "mul"s can be issued every 3 cycles. This suggests 3 is a limit on
69 1.1 mrg the speed of the multiplication routines. The documentation shows mul
70 1.1 mrg executing in IEU0 (or maybe in IEU0 and IEU1 together), so it might be that,
71 1.1 mrg to get near 3 cycles code has to be arranged so that nothing else is issued
72 1.1 mrg to IEU0. A busy IEU0 could explain why some code takes 4 cycles and other
73 1.1 mrg apparently equivalent code takes 5.
74 1.1 mrg
75 1.1 mrg
76 1.1 mrg
77 1.1 mrg OPTIMIZATIONS
78 1.1 mrg
79 1.1 mrg Unrolled loops are used to reduce looping overhead. The unrolling is
80 1.1 mrg configurable up to 32 limbs/loop for most routines and up to 64 for some.
81 1.1 mrg The K7 has 64k L1 code cache so quite big unrolling is allowable.
82 1.1 mrg
83 1.1 mrg Computed jumps into the unrolling are used to handle sizes not a multiple of
84 1.1 mrg the unrolling. An attractive feature of this is that times increase
85 1.1 mrg smoothly with operand size, but it may be that some routines should just
86 1.1 mrg have simple loops to finish up, especially when PIC adds between 2 and 16
87 1.1 mrg cycles to get %eip.
88 1.1 mrg
89 1.1 mrg Position independent code is implemented using a call to get %eip for the
90 1.1 mrg computed jumps and a ret is always done, rather than an addl $4,%esp or a
91 1.1 mrg popl, so the CPU return address branch prediction stack stays synchronised
92 1.1 mrg with the actual stack in memory.
93 1.1 mrg
94 1.1 mrg Branch prediction, in absence of any history, will guess forward jumps are
95 1.1 mrg not taken and backward jumps are taken. Where possible it's arranged that
96 1.1 mrg the less likely or less important case is under a taken forward jump.
97 1.1 mrg
98 1.1 mrg
99 1.1 mrg
100 1.1 mrg CODING
101 1.1 mrg
102 1.1 mrg Instructions in general code have been shown grouped if they can execute
103 1.1 mrg together, which means up to three direct-path instructions which have no
104 1.1 mrg successive dependencies. K7 always decodes three and has out-of-order
105 1.1 mrg execution, but the groupings show what slots might be available and what
106 1.1 mrg dependency chains exist.
107 1.1 mrg
108 1.1 mrg When there's vector-path instructions an effort is made to get triplets of
109 1.1 mrg direct-path instructions in between them, even if there's dependencies,
110 1.1 mrg since this maximizes decoding throughput and might save a cycle or two if
111 1.1 mrg decoding is the limiting factor.
112 1.1 mrg
113 1.1 mrg
114 1.1 mrg
115 1.1 mrg INSTRUCTIONS
116 1.1 mrg
117 1.1 mrg adcl direct
118 1.1 mrg divl 39 cycles back-to-back
119 1.1 mrg lodsl,etc vector
120 1.1 mrg loop 1 cycle vector (decl/jnz opens up one decode slot)
121 1.1 mrg movd reg vector
122 1.1 mrg movd mem direct
123 1.1 mrg mull issue every 3 cycles, latency 4 cycles low word, 6 cycles high word
124 1.1 mrg popl vector (use movl for more than one pop)
125 1.1 mrg pushl direct, will pair with a load
126 1.1 mrg shrdl %cl vector, 3 cycles, seems to be 3 decode too
127 1.1 mrg xorl r,r false read dependency recognised
128 1.1 mrg
129 1.1 mrg
130 1.1 mrg
131 1.1 mrg REFERENCES
132 1.1 mrg
133 1.1 mrg "AMD Athlon Processor X86 Code Optimization Guide", AMD publication number
134 1.1 mrg 22007, revision K, February 2002. Available on-line,
135 1.1 mrg
136 1.1 mrg http://www.amd.com/us-en/assets/content_type/white_papers_and_tech_docs/22007.pdf
137 1.1 mrg
138 1.1 mrg "3DNow Technology Manual", AMD publication number 21928G/0-March 2000.
139 1.1 mrg This describes the femms and prefetch instructions. Available on-line,
140 1.1 mrg
141 1.1 mrg http://www.amd.com/us-en/assets/content_type/white_papers_and_tech_docs/21928.pdf
142 1.1 mrg
143 1.1 mrg "AMD Extensions to the 3DNow and MMX Instruction Sets Manual", AMD
144 1.1 mrg publication number 22466, revision D, March 2000. This describes
145 1.1 mrg instructions added in the Athlon processor, such as pswapd and the extra
146 1.1 mrg prefetch forms. Available on-line,
147 1.1 mrg
148 1.1 mrg http://www.amd.com/us-en/assets/content_type/white_papers_and_tech_docs/22466.pdf
149 1.1 mrg
150 1.1 mrg "3DNow Instruction Porting Guide", AMD publication number 22621, revision B,
151 1.1 mrg August 1999. This has some notes on general Athlon optimizations as well as
152 1.1 mrg 3DNow. Available on-line,
153 1.1 mrg
154 1.1 mrg http://www.amd.com/us-en/assets/content_type/white_papers_and_tech_docs/22621.pdf
155 1.1 mrg
156 1.1 mrg
157 1.1 mrg
158 1.1 mrg
159 1.1 mrg ----------------
160 1.1 mrg Local variables:
161 1.1 mrg mode: text
162 1.1 mrg fill-column: 76
163 1.1 mrg End:
164