FFmpeg

mirror of https://github.com/FFmpeg/FFmpeg.git synced 2025-02-14 22:22:59 +02:00

History

Martin Storsjö 9532a7d4d0 aarch64: vp9itxfm: Do separate functions for half/quarter idct16 and idct32

This work is sponsored by, and copyright, Google.

This avoids loading and calculating coefficients that we know will
be zero, and avoids filling the temp buffer with zeros in places
where we know the second pass won't read.

This gives a pretty substantial speedup for the smaller subpartitions.

The code size increases from 14740 bytes to 24292 bytes.

The idct16/32_end macros are moved above the individual functions; the
instructions themselves are unchanged, but since new functions are added
at the same place where the code is moved from, the diff looks rather
messy.

Before:
vp9_inv_dct_dct_16x16_sub1_add_neon:     236.7
vp9_inv_dct_dct_16x16_sub2_add_neon:    1051.0
vp9_inv_dct_dct_16x16_sub4_add_neon:    1051.0
vp9_inv_dct_dct_16x16_sub8_add_neon:    1051.0
vp9_inv_dct_dct_16x16_sub12_add_neon:   1387.4
vp9_inv_dct_dct_16x16_sub16_add_neon:   1387.6
vp9_inv_dct_dct_32x32_sub1_add_neon:     554.1
vp9_inv_dct_dct_32x32_sub2_add_neon:    5198.5
vp9_inv_dct_dct_32x32_sub4_add_neon:    5198.6
vp9_inv_dct_dct_32x32_sub8_add_neon:    5196.3
vp9_inv_dct_dct_32x32_sub12_add_neon:   6183.4
vp9_inv_dct_dct_32x32_sub16_add_neon:   6174.3
vp9_inv_dct_dct_32x32_sub20_add_neon:   7151.4
vp9_inv_dct_dct_32x32_sub24_add_neon:   7145.3
vp9_inv_dct_dct_32x32_sub28_add_neon:   8119.3
vp9_inv_dct_dct_32x32_sub32_add_neon:   8118.7

After:
vp9_inv_dct_dct_16x16_sub1_add_neon:     236.7
vp9_inv_dct_dct_16x16_sub2_add_neon:     640.8
vp9_inv_dct_dct_16x16_sub4_add_neon:     639.0
vp9_inv_dct_dct_16x16_sub8_add_neon:     842.0
vp9_inv_dct_dct_16x16_sub12_add_neon:   1388.3
vp9_inv_dct_dct_16x16_sub16_add_neon:   1389.3
vp9_inv_dct_dct_32x32_sub1_add_neon:     554.1
vp9_inv_dct_dct_32x32_sub2_add_neon:    3685.5
vp9_inv_dct_dct_32x32_sub4_add_neon:    3685.1
vp9_inv_dct_dct_32x32_sub8_add_neon:    3684.4
vp9_inv_dct_dct_32x32_sub12_add_neon:   5312.2
vp9_inv_dct_dct_32x32_sub16_add_neon:   5315.4
vp9_inv_dct_dct_32x32_sub20_add_neon:   7154.9
vp9_inv_dct_dct_32x32_sub24_add_neon:   7154.5
vp9_inv_dct_dct_32x32_sub28_add_neon:   8126.6
vp9_inv_dct_dct_32x32_sub32_add_neon:   8127.2

This is cherrypicked from libav commit
a63da4511d0fee66695ff4afd264ba1dbf1e812d.

Signed-off-by: Martin Storsjö <martin@martin.st>

2017-03-11 13:14:25 +02:00

asm-offsets.h

…

cabac.h

…

fft_init_aarch64.c

…

fft_neon.S

…

fmtconvert_init.c

…

fmtconvert_neon.S

…

h264chroma_init_aarch64.c

…

h264cmc_neon.S

…

h264dsp_init_aarch64.c

…

h264dsp_neon.S

…

h264idct_neon.S

aarch64: h264idct: Use the offset parameter to movrel

2016-12-08 18:11:07 +01:00

h264pred_init.c

…

h264pred_neon.S

…

h264qpel_init_aarch64.c

…

h264qpel_neon.S

…

hpeldsp_init_aarch64.c

…

hpeldsp_neon.S

…

Makefile

aarch64: Add NEON optimizations for 10 and 12 bit vp9 loop filter

2017-01-24 22:36:11 +02:00

mdct_neon.S

…

mpegaudiodsp_init.c

…

mpegaudiodsp_neon.S

…

neon.S

…

neontest.c

avcodec: fix arguments on xmm/neon clobber test wrappers

2016-10-02 02:15:47 -03:00

rv40dsp_init_aarch64.c

…

synth_filter_init.c

…

synth_filter_neon.S

…

vc1dsp_init_aarch64.c

…

videodsp_init.c

…

videodsp.S

…

vorbisdsp_init.c

…

vorbisdsp_neon.S

…

vp9dsp_init_10bpp_aarch64.c

aarch64: Add NEON optimizations for 10 and 12 bit vp9 MC

2017-01-24 22:36:05 +02:00

vp9dsp_init_12bpp_aarch64.c

aarch64: Add NEON optimizations for 10 and 12 bit vp9 MC

2017-01-24 22:36:05 +02:00

vp9dsp_init_16bpp_aarch64_template.c

aarch64: Add NEON optimizations for 10 and 12 bit vp9 loop filter

2017-01-24 22:36:11 +02:00

vp9dsp_init_aarch64.c

aarch64: Add NEON optimizations for 10 and 12 bit vp9 MC

2017-01-24 22:36:05 +02:00

vp9dsp_init.h

aarch64: Add NEON optimizations for 10 and 12 bit vp9 MC

2017-01-24 22:36:05 +02:00

vp9itxfm_16bpp_neon.S

aarch64: Add NEON optimizations for 10 and 12 bit vp9 itxfm

2017-01-24 22:36:08 +02:00

vp9itxfm_neon.S

aarch64: vp9itxfm: Do separate functions for half/quarter idct16 and idct32

2017-03-11 13:14:25 +02:00

vp9lpf_16bpp_neon.S

aarch64: Add NEON optimizations for 10 and 12 bit vp9 loop filter

2017-01-24 22:36:11 +02:00

vp9lpf_neon.S

aarch64: vp9: loop filter: replace 'orr; cbn?z' with 'adds; b.{eq,ne};

2017-01-14 21:13:10 +01:00

vp9mc_16bpp_neon.S

aarch64: Add NEON optimizations for 10 and 12 bit vp9 MC

2017-01-24 22:36:05 +02:00

vp9mc_neon.S

aarch64: vp9mc: Fix a comment to refer to a register with the right name

2017-01-14 21:13:43 +01:00