Commit 5686c8c
authored
Improve element-wise tensor operation performance (#454)
transform_tensor previously used fplus::transform_convert, which calls
reserve() and writes via a back_insert_iterator. Because push_back
updates the vector's internal end pointer on every write, the compiler
cannot auto-vectorize the loop.
Switching to resize() with a plain random-access iterator allows GCC to
emit SIMD instructions. Additionally, relu_layer now takes a fast path
for the common standard ReLU case, using a simpler single-expression
lambda that is easier to optimize than the general three-branch version.
Combined effect: ~5% faster forward pass on VGG19.1 parent 75f5610 commit 5686c8c
2 files changed
Lines changed: 14 additions & 4 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
9 | 9 | | |
10 | 10 | | |
11 | 11 | | |
| 12 | + | |
12 | 13 | | |
13 | 14 | | |
14 | 15 | | |
| |||
30 | 31 | | |
31 | 32 | | |
32 | 33 | | |
33 | | - | |
| 34 | + | |
| 35 | + | |
| 36 | + | |
| 37 | + | |
| 38 | + | |
| 39 | + | |
| 40 | + | |
34 | 41 | | |
35 | 42 | | |
36 | 43 | | |
37 | 44 | | |
38 | 45 | | |
39 | | - | |
40 | | - | |
| 46 | + | |
| 47 | + | |
41 | 48 | | |
42 | 49 | | |
43 | 50 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
190 | 190 | | |
191 | 191 | | |
192 | 192 | | |
193 | | - | |
| 193 | + | |
| 194 | + | |
| 195 | + | |
| 196 | + | |
194 | 197 | | |
195 | 198 | | |
196 | 199 | | |
| |||
0 commit comments