FFmpeg

mirror of https://github.com/FFmpeg/FFmpeg.git synced 2024-12-12 19:18:44 +02:00

Author	SHA1	Message	Date
Hendrik Leppkes	fc7e02f0ff	dcadsp: fix SSE code to not use SSE2 instructions. movq from SSE register to memory is an SSE2 instruction. Instead, use SSE movlps, which does the same thing. Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-04-06 18:31:22 +02:00
James Almer	a1ac12bddd	x86/dcadsp: add ff_dca_lfe_fir0_fma3 ~10% faster than the SSE version. Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-04-05 13:55:59 +02:00
James Almer	7d2116dd09	x86/synth_filter: compile avx and fma3 functions unconditionally Fixes compilation failures with "--disable-{avx,fma3} --disable-optimizations" Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-04-05 05:15:27 +02:00
Michael Niedermayer	51fd962c0b	Merge commit 'c74b86699c86bdf62e8570f41d8a38be5710baa3' * commit 'c74b86699c86bdf62e8570f41d8a38be5710baa3': x86/synth_filter: add synth_filter_fma3 x86/synth_filter: add synth_filter_avx x86/synth_filter: add synth_filter_sse Conflicts: libavcodec/x86/dcadsp.asm libavcodec/x86/dcadsp_init.c See: `6467209836` See: `68c3ed936a` See: `7fd64e3e36` See: `aa1f38015c` See: `dfd865e51b` Merged-by: Michael Niedermayer <michaelni@gmx.at>	2014-04-04 23:40:08 +02:00
Christophe Gisquet	dfd865e51b	x86/synth_filter: remove the main loop when it's not needed Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-04-04 22:35:45 +02:00
James Almer	c74b86699c	x86/synth_filter: add synth_filter_fma3 Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Anton Khirnov <anton@khirnov.net>	2014-04-04 17:40:51 +02:00
James Almer	81e02fae6e	x86/synth_filter: add synth_filter_avx Sandy Bridge Win64: 180 cycles in ff_synth_filter_inner_sse2 150 cycles in ff_synth_filter_inner_avx Also switch some instructions to a three operand format to avoid assembly errors with Yasm 1.1.0 or older. Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Anton Khirnov <anton@khirnov.net>	2014-04-04 17:40:51 +02:00
James Almer	2025d8026f	x86/synth_filter: add synth_filter_sse Build only on x86_32 targets. Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Anton Khirnov <anton@khirnov.net>	2014-04-04 17:40:51 +02:00
James Almer	aa1f38015c	x86/synth_filter: improve FMA version Replace mulps+subps with fnmaddps, resulting in two less instructions inside the inner loops. About 1% faster FMA3 performance. Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-03-17 21:04:15 +01:00
James Almer	7fd64e3e36	x86/synth_filter: add synth_filter_fma3 Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-03-05 01:58:16 +01:00
James Almer	884e085d1e	x86/synth_filter: Revert the switch to float ops with SSE2 This reverts the changes `6467209836` and `68c3ed936a` did to the SSE2 version, which generated a hit of about 5 cycles. Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-03-02 11:58:10 +01:00
James Almer	68c3ed936a	x86/synth_filter: add synth_filter_avx Sandy Bridge Win64: 180 cycles on ff_synth_filter_inner_sse2 150 cycles on ff_synth_filter_inner_avx Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-03-02 01:00:55 +01:00
James Almer	6467209836	x86/synth_filter: add synth_filter_sse Build only on x86_32 targets. Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-03-01 15:32:40 +01:00
Michael Niedermayer	fb3c33f3cd	Merge commit '4cb6964244fd6c099383d8b7e99731e72cc844b9' * commit '4cb6964244fd6c099383d8b7e99731e72cc844b9': dcadec: simplify decoding of VQ high frequencies Conflicts: configure libavcodec/dcadec.c Merged-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-28 21:41:19 +01:00
Michael Niedermayer	baf3adc621	Merge commit '08e3ea60ff4059341b74be04a428a38f7c3630b0' * commit '08e3ea60ff4059341b74be04a428a38f7c3630b0': x86: synth filter float: implement SSE2 version Conflicts: libavcodec/x86/dcadsp.asm libavcodec/x86/dcadsp_init.c See: `2cdbcc0048` Merged-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-28 20:38:39 +01:00
Christophe Gisquet	2cdbcc0048	x86: synth filter float: implement SSE2 version Timings for Arrandale: C SSE win32: 2108 334 win64: 1152 322 Factorizing the inner loop with a call/jmp is a >15 cycles cost, even with the jmp destination being aligned. Unrolling for ARCH_X86_64 is a 20 cycles gain. Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-28 20:34:40 +01:00
Michael Niedermayer	e346a59383	Merge commit 'ad507d7907457e678900bac132122ba7be4644cb' * commit 'ad507d7907457e678900bac132122ba7be4644cb': x86: dcadsp: implement SSE lfe_dir Conflicts: libavcodec/x86/dcadsp.asm See: `169243112c` Merged-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-28 19:22:00 +01:00
Christophe Gisquet	169243112c	x86: dcadsp: implement SSE lfe_dir Results for Arrandale/Windows: 32: 1670 -> 316 64: 728 -> 298 Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-28 19:20:03 +01:00
Christophe Gisquet	4cb6964244	dcadec: simplify decoding of VQ high frequencies The vector dequantization has a test in a loop preventing effective SIMD implementation. By moving it out of the loop, this loop can be DSPized. Therefore, modify the current DSP implementation. In particular, the DSP implementation no longer has to handle null loop sizes. The decode_hf implementations have following timings: For x86 Arrandale: C SSE SSE2 SSE4 win32: 260 162 119 104 win64: 242 N/A 89 72 The arm NEON optimizations follow in a later patch as external asm. The now unused check for the y modifier in arm inline asm is removed from configure.	2014-02-28 13:03:22 +01:00
Christophe Gisquet	08e3ea60ff	x86: synth filter float: implement SSE2 version Timings for Arrandale: C SSE win32: 2108 334 win64: 1152 322 Factorizing the inner loop with a call/jmp is a >15 cycles cost, even with the jmp destination being aligned. Unrolling for ARCH_X86_64 is a 20 cycles gain. Signed-off-by: Janne Grunau <janne-libav@jannau.net>	2014-02-28 13:00:48 +01:00
Christophe Gisquet	ad507d7907	x86: dcadsp: implement SSE lfe_dir Results for Arrandale/Windows: 32: 1670 -> 316 64: 728 -> 298 Signed-off-by: Janne Grunau <janne-libav@jannau.net>	2014-02-28 13:00:47 +01:00
Michael Niedermayer	82ae8a44e6	Merge commit '5b59a9fc6152169599561f04b4f66370edda5c9c' * commit '5b59a9fc6152169599561f04b4f66370edda5c9c': x86: dcadsp: implement int8x8_fmul_int32 Merged-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-08 01:20:33 +01:00
Christophe Gisquet	5b59a9fc61	x86: dcadsp: implement int8x8_fmul_int32 For the callable function (as opposed to the inline one): C SSE SSE2 SSE4 Win32: 47 42 29 26 Win64: 30 33 25 23 The SSE version is neither compiled nor set for ARCH_X86_64, as the inlinable function takes over. Signed-off-by: Janne Grunau <janne-libav@jannau.net>	2014-02-07 22:52:40 +01:00

23 Commits