FFmpeg

mirror of https://github.com/FFmpeg/FFmpeg.git synced 2024-12-18 03:19:31 +02:00

Author	SHA1	Message	Date
Diego Biurrun	994c4bc107	x86util: Port all macros to cpuflags Also do some small cosmetic changes: Drop pointless _MMX suffix from ABSD2 macro name, drop pointless check for MMX support, we always assume MMX is available in our SIMD code, fix spelling.	2017-03-14 17:23:32 +01:00
Diego Biurrun	39e208f4d4	build: Generalize yasm/nasm-related variable names None of them are specific to the YASM assembler.	2017-03-01 10:18:15 +01:00
James Darnley	5336887867	avcodec/h264: sse2, avx h luma mbaff deblock/loop filter x86-64 only Yorkfield: - sse2: ~2.17x (434 vs. 200 cycles) Nehalem: - sse2: ~2.94x (409 vs. 139 cycles) Skylake: - sse2: ~3.10x (370 vs. 119 cycles) - avx: ~3.29x (370 vs. 112 cycles)	2017-02-18 20:26:52 +01:00
James Darnley	7627df15d4	x86util: import MOVHL macro Originally committed to x264 in 1637239a by Henrik Gramner who has agreed to re-license it as LGPL. Original commit message follows. x86: Avoid some bypass delays and false dependencies A bypass delay of 1-3 clock cycles may occur on some CPUs when transitioning between int and float domains, so try to avoid that if possible.	2017-02-18 20:26:51 +01:00
James Darnley	9d815b7424	avcodec/x86: deduplicate PASS8ROWS macro	2017-02-18 20:26:49 +01:00
Diego Biurrun	7abdd026df	asm: Consistently uppercase SECTION markers	2017-02-03 11:37:53 +01:00
James Almer	8d5df204d0	Merge commit '8e9cd81d291b1010c625b2766058aadf4affb537' * commit '8e9cd81d291b1010c625b2766058aadf4affb537': x86: cpu: Detect Conroe CPUs and their slow shuffle unit Merged-by: James Almer <jamrial@gmail.com>	2017-01-31 15:20:54 -03:00
James Almer	2eab48177d	Merge commit '7d7355aa92bb36ca0765c49a569a999bcb96f332' * commit '7d7355aa92bb36ca0765c49a569a999bcb96f332': x86: Add SSSE3_SLOW CPU flag and related convenience macros Merged-by: James Almer <jamrial@gmail.com>	2017-01-31 15:17:19 -03:00
Henrik Gramner	cd09e3b349	x86inc: Avoid using eax/rax for storing the stack pointer When allocating stack space with an alignment requirement that is larger than the current stack alignment we need to store a copy of the original stack pointer in order to be able to restore it later. If we chose to use another register for this purpose we should not pick eax/rax since it can be overwritten as a return value.	2017-01-09 16:00:29 +01:00
Henrik Gramner	3cba1ad76d	x86inc: Avoid using eax/rax for storing the stack pointer When allocating stack space with an alignment requirement that is larger than the current stack alignment we need to store a copy of the original stack pointer in order to be able to restore it later. If we chose to use another register for this purpose we should not pick eax/rax since it can be overwritten as a return value. Signed-off-by: Anton Khirnov <anton@khirnov.net>	2017-01-09 13:21:12 +01:00
Diego Biurrun	99434f4df8	float_dsp: Have implementation match function pointer prototype libavutil/x86/float_dsp_init.c(144) : warning C4028: formal parameter 1 different from declaration libavutil/x86/float_dsp_init.c(144) : warning C4028: formal parameter 2 different from declaration	2016-11-03 17:43:55 +01:00
Michael Niedermayer	051517648b	avutil/x86/emms: Document the emms_c() vs alloc/free relation. Reviewed-by: Andreas Cadhalpun <andreas.cadhalpun@googlemail.com> Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>	2016-10-23 13:02:37 +02:00
Diego Biurrun	7911186ed6	emms: Give apriv_emms_yasm() a more general name	2016-10-18 13:09:09 +02:00
Diego Biurrun	6be7944ee2	x86: Add missing colons after assembly labels This fixes many warnings of the sort warning: label alone on a line without a colon might be in error	2016-10-17 16:31:26 +02:00
Alexandra Hájková	07e1f99a1b	x86util: Document SBUTTERFLY macro Signed-off-by: Luca Barbato <lu_zero@gentoo.org>	2016-09-19 10:02:43 +02:00
Anton Khirnov	d7bc52bf45	imgutils: add a function for copying image data from GPU mapped memory See https://software.intel.com/en-us/articles/copying-accelerated-video-decode-frame-buffers	2016-08-31 08:15:47 +02:00
Fiona Glaser	8e9cd81d29	x86: cpu: Detect Conroe CPUs and their slow shuffle unit	2016-07-20 18:43:28 +02:00
Diego Biurrun	7d7355aa92	x86: Add SSSE3_SLOW CPU flag and related convenience macros	2016-07-20 18:43:28 +02:00
James Almer	fd5e6a095f	x86util: Extend SPLATW for avx2 Integration to Libav by Josh de Kock <josh@itanimul.li>. Signed-off-by: Alexandra Hájková <alexandra@khirnov.net>	2016-07-18 15:27:13 +02:00
Ronald S. Bultje	f0a2b6249b	vp9: add 16x16 idct avx2 (8-bit). checkasm --bench, 10k runs, for *_add_${bpc}_${sub_idct}_${opt}, shows that it's about 1.65x as fast as the AVX version for the full IDCT, and similar speedups for the sub-IDCTs: nop: 24.6 vp9_inv_dct_dct_16x16_add_8_1_c: 6444.8 vp9_inv_dct_dct_16x16_add_8_1_sse2: 638.6 vp9_inv_dct_dct_16x16_add_8_1_ssse3: 484.4 vp9_inv_dct_dct_16x16_add_8_1_avx: 661.2 vp9_inv_dct_dct_16x16_add_8_1_avx2: 311.5 vp9_inv_dct_dct_16x16_add_8_2_c: 6665.7 vp9_inv_dct_dct_16x16_add_8_2_sse2: 646.9 vp9_inv_dct_dct_16x16_add_8_2_ssse3: 455.2 vp9_inv_dct_dct_16x16_add_8_2_avx: 521.9 vp9_inv_dct_dct_16x16_add_8_2_avx2: 304.3 vp9_inv_dct_dct_16x16_add_8_4_c: 7022.7 vp9_inv_dct_dct_16x16_add_8_4_sse2: 647.4 vp9_inv_dct_dct_16x16_add_8_4_ssse3: 467.1 vp9_inv_dct_dct_16x16_add_8_4_avx: 446.1 vp9_inv_dct_dct_16x16_add_8_4_avx2: 297.0 vp9_inv_dct_dct_16x16_add_8_8_c: 6800.4 vp9_inv_dct_dct_16x16_add_8_8_sse2: 598.6 vp9_inv_dct_dct_16x16_add_8_8_ssse3: 465.7 vp9_inv_dct_dct_16x16_add_8_8_avx: 440.9 vp9_inv_dct_dct_16x16_add_8_8_avx2: 290.2 vp9_inv_dct_dct_16x16_add_8_16_c: 6626.6 vp9_inv_dct_dct_16x16_add_8_16_sse2: 599.5 vp9_inv_dct_dct_16x16_add_8_16_ssse3: 475.0 vp9_inv_dct_dct_16x16_add_8_16_avx: 469.9 vp9_inv_dct_dct_16x16_add_8_16_avx2: 286.4	2016-07-11 10:14:58 -04:00
Matthieu Bouron	9eb3da2f99	asm: FF_-prefix internal macros used in inline assembly See merge commit '39d6d3618d48625decaff7d9bdbb45b44ef2a805'.	2016-06-27 17:21:18 +02:00
Clément Bœsch	8ef57a0d61	Merge commit '41ed7ab45fc693f7d7fc35664c0233f4c32d69bb' * commit '41ed7ab45fc693f7d7fc35664c0233f4c32d69bb': cosmetics: Fix spelling mistakes Merged-by: Clément Bœsch <u@pkh.me>	2016-06-21 21:55:34 +02:00
Matt Oliver	5ca44ebd99	lavu/intmath.h: fix compilation with msvc10. Signed-off-by: Matt Oliver <protogonoi@gmail.com>	2016-06-13 13:49:24 +10:00
James Almer	172af20852	x86/showcqt: use three operand format for some instructions Fixes failures with yasm 1.1.0 and older Signed-off-by: James Almer <jamrial@gmail.com>	2016-06-08 19:37:08 -03:00
James Almer	99b899483e	avutil/x86util: move haddps sse emulation from showcqt Signed-off-by: James Almer <jamrial@gmail.com>	2016-06-08 14:18:00 -03:00
Diego Biurrun	1e9c5bf4c1	asm: FF_-prefix internal macros used in inline assembly These warnings conflict with system macros on Solaris, producing truckloads of warnings about macro redefinition.	2016-05-28 19:18:26 +02:00
Anton Mitrofanov	2fb1d17a5a	x86inc: Enable AVX emulation in additional cases Allows emulation to work when dst is equal to src2 as long as the instruction is commutative, e.g. `addps m0, m1, m0`. Signed-off-by: Anton Khirnov <anton@khirnov.net>	2016-05-16 10:31:24 +02:00
Anton Mitrofanov	300fb0df84	x86inc: Improve handling of %ifid with multi-token parameters The yasm/nasm preprocessor only checks the first token, which means that parameters such as `dword [rax]` are treated as identifiers, which is generally not what we want. Signed-off-by: Anton Khirnov <anton@khirnov.net>	2016-05-16 10:31:20 +02:00
Anton Mitrofanov	8d02579fae	x86inc: Fix AVX emulation of some instructions Signed-off-by: Anton Khirnov <anton@khirnov.net>	2016-05-16 10:31:17 +02:00
Henrik Gramner	ba3eb745cc	x86inc: Fix AVX emulation of scalar float instructions Those instructions are not commutative since they only change the first element in the vector and leave the rest unmodified. Signed-off-by: Anton Khirnov <anton@khirnov.net>	2016-05-16 10:31:13 +02:00
Vittorio Giovara	41ed7ab45f	cosmetics: Fix spelling mistakes Signed-off-by: Diego Biurrun <diego@biurrun.de>	2016-05-04 18:16:21 +02:00
Anton Mitrofanov	e428f3b30c	x86inc: Enable AVX emulation in additional cases Allows emulation to work when dst is equal to src2 as long as the instruction is commutative, e.g. `addps m0, m1, m0`.	2016-04-20 19:16:22 +02:00
Anton Mitrofanov	4bd5583ace	x86inc: Improve handling of %ifid with multi-token parameters The yasm/nasm preprocessor only checks the first token, which means that parameters such as `dword [rax]` are treated as identifiers, which is generally not what we want.	2016-04-20 19:16:22 +02:00
Anton Mitrofanov	42be240ad6	x86inc: Fix AVX emulation of some instructions	2016-04-20 19:16:22 +02:00
Henrik Gramner	8dd3ee9ddd	x86inc: Fix AVX emulation of scalar float instructions Those instructions are not commutative since they only change the first element in the vector and leave the rest unmodified.	2016-04-20 19:16:22 +02:00
James Almer	70d685a77f	x86: use the new helper macros where useful Reviewed-by: Michael Niedermayer <michael@niedermayer.cc> Signed-off-by: James Almer <jamrial@gmail.com>	2016-02-14 20:00:21 -03:00
James Almer	73a4589d4b	x86: add some more helper macros to check for slow cpuflags Reviewed-by: Michael Niedermayer <michael@niedermayer.cc> Signed-off-by: James Almer <jamrial@gmail.com>	2016-02-14 20:00:17 -03:00
James Almer	be22bd32fe	x86/cpu: set avxslow cpuflag on btver2 CPUs They are also slow when using 256 bit wide registers Reviewed-by: Hendrik Leppkes <h.leppkes@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	2016-02-07 16:39:21 -03:00
James Almer	b3b0ecee15	x86/emms: empty the mmx state unconditionally on supported targets Reviewed-by: Michael Niedermayer <michael@niedermayer.cc> Signed-off-by: James Almer <jamrial@gmail.com>	2016-02-04 01:49:01 -03:00
Timothy Gu	44304ae322	all: Add missing header guards	2016-01-28 19:49:48 -08:00
James Almer	b624f0660b	x86: Add ymm_reg struct Needed to declare 32-byte long constants Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Luca Barbato <lu_zero@gentoo.org>	2016-01-28 00:41:19 +01:00
Geza Lore	cc602061ee	x86inc: Add debug symbols indicating sizes of compiled functions Some debuggers/profilers use this metadata to determine which function a given instruction is in; without it they get can confused by local labels (if you haven't stripped those). On the other hand, some tools are still confused even with this metadata. e.g. this fixes `gdb`, but not `perf`. Currently only implemented for ELF. Signed-off-by: Anton Khirnov <anton@khirnov.net>	2016-01-23 20:46:28 +01:00
Henrik Gramner	002c47798d	x86inc: Avoid creating unnecessary local labels The REP_RET workaround is only needed on old AMD cpus, and the labels clutter up the symbol table and confuse debugging/profiling tools, so use EQU to create SHN_ABS symbols instead of creating local labels. Furthermore, skip the workaround completely in functions that definitely won't run on such cpus. Note that EQU is just creating a local label when using nasm instead of yasm. This is probably a bug, but at least it doesn't break anything. Signed-off-by: Anton Khirnov <anton@khirnov.net>	2016-01-23 20:44:25 +01:00
Henrik Gramner	fd6ecac38e	x86inc: Simplify AUTO_REP_RET cpuflags is never undefined any more, it's set to 0 instead. Also fix an incorrect comment. Signed-off-by: Anton Khirnov <anton@khirnov.net>	2016-01-23 20:43:39 +01:00
Henrik Gramner	5ca8e195e5	x86inc: Use more consistent indentation Signed-off-by: Anton Khirnov <anton@khirnov.net>	2016-01-23 20:42:59 +01:00
Henrik Gramner	91ed050f42	x86inc: Preserve arguments when allocating stack space When allocating stack space with a larger alignment than the known stack alignment a temporary register is used for storing the stack pointer. Ensure that this isn't one of the registers used for passing arguments. Signed-off-by: Anton Khirnov <anton@khirnov.net>	2016-01-23 20:41:59 +01:00
Henrik Gramner	715eb7ca24	x86inc: Improve FMA instruction handling * Correctly handle FMA instructions with memory operands. * Print a warning if FMA instructions are used without the correct cpuflag. * Simplify the instantiation code. * Clarify documentation. Only the last operand in FMA3 instructions can be a memory operand. When converting FMA4 instructions to FMA3 instructions we can utilize the fact that multiply is a commutative operation and reorder operands if necessary to ensure that a memory operand is used only as the last operand. Signed-off-by: Anton Khirnov <anton@khirnov.net>	2016-01-23 20:30:30 +01:00
Henrik Gramner	f60f06d989	x86inc: Be more verbose in assertion failures Signed-off-by: Anton Khirnov <anton@khirnov.net>	2016-01-23 20:30:07 +01:00
Henrik Gramner	7adcd4e841	x86inc: Make cpuflag() and notcpuflag() return 0 or 1 Makes it possible to use them in arithmetic expressions. Signed-off-by: Anton Khirnov <anton@khirnov.net>	2016-01-23 20:19:19 +01:00
Geza Lore	d39c229e54	x86inc: Add debug symbols indicating sizes of compiled functions Some debuggers/profilers use this metadata to determine which function a given instruction is in; without it they get can confused by local labels (if you haven't stripped those). On the other hand, some tools are still confused even with this metadata. e.g. this fixes `gdb`, but not `perf`. Currently only implemented for ELF.	2016-01-21 23:19:46 +01:00
Henrik Gramner	d3662777e0	x86inc: Avoid creating unnecessary local labels The REP_RET workaround is only needed on old AMD cpus, and the labels clutter up the symbol table and confuse debugging/profiling tools, so use EQU to create SHN_ABS symbols instead of creating local labels. Furthermore, skip the workaround completely in functions that definitely won't run on such cpus. Note that EQU is just creating a local label when using nasm instead of yasm. This is probably a bug, but at least it doesn't break anything.	2016-01-21 23:19:46 +01:00
Henrik Gramner	87b587d4fe	x86inc: Simplify AUTO_REP_RET cpuflags is never undefined any more, it's set to 0 instead. Also fix an incorrect comment.	2016-01-21 23:19:46 +01:00
Henrik Gramner	2d60b18cf0	x86inc: Use more consistent indentation	2016-01-21 23:19:46 +01:00
Henrik Gramner	dfe771dc5a	x86inc: Preserve arguments when allocating stack space When allocating stack space with a larger alignment than the known stack alignment a temporary register is used for storing the stack pointer. Ensure that this isn't one of the registers used for passing arguments.	2016-01-21 23:19:46 +01:00
Henrik Gramner	b1496008ee	x86inc: Improve FMA instruction handling * Correctly handle FMA instructions with memory operands. * Print a warning if FMA instructions are used without the correct cpuflag. * Simplify the instantiation code. * Clarify documentation. Only the last operand in FMA3 instructions can be a memory operand. When converting FMA4 instructions to FMA3 instructions we can utilize the fact that multiply is a commutative operation and reorder operands if necessary to ensure that a memory operand is used only as the last operand.	2016-01-21 23:19:46 +01:00
Henrik Gramner	6cbd0fdf28	x86inc: Be more verbose in assertion failures	2016-01-21 23:19:46 +01:00
James Almer	a2e1b66460	x86/intmath: disable sse av_clip functions when using ICC It seems to miscompile them Should fix fate-ra-288 and fate-twinvq Reviewed-by: Michael Niedermayer <michael@niedermayer.cc> Signed-off-by: James Almer <jamrial@gmail.com>	2016-01-21 16:50:51 -03:00
James Almer	dee579ffcd	x86/fixed_dsp: add ff_butterflies_fixed_sse2 Reviewed-by: Paul B Mahol <onemda@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	2016-01-16 21:09:38 -03:00
Ganesh Ajjanagadde	5989add4ab	lavu/x86/lls: add fma3 optimizations for update_lls This improves accuracy (very slightly) and speed for processors having fma3. Sample benchmark (fate flac-16-lpc-cholesky, Haswell): old: 5993610 decicycles in ff_lpc_calc_coefs, 64 runs, 0 skips 5951528 decicycles in ff_lpc_calc_coefs, 128 runs, 0 skips new: 5252410 decicycles in ff_lpc_calc_coefs, 64 runs, 0 skips 5232869 decicycles in ff_lpc_calc_coefs, 128 runs, 0 skips Tested with FATE and --disable-fma3, also examined contents of lavu/lls-test. Reviewed-by: James Almer <jamrial@gmail.com> Reviewed-by: Henrik Gramner <henrik@gramner.com> Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>	2016-01-15 16:46:13 -05:00
James Almer	36778627e2	x86/intmath: add missing early clobber to output operands Signed-off-by: James Almer <jamrial@gmail.com>	2016-01-15 13:32:58 -03:00
James Almer	dc79824deb	x86/float_dsp: zero extend offset from ff_scalarproduct_float_sse Reviewed-by: Christophe Gisquet <christophe.gisquet@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	2016-01-08 16:14:32 -03:00
James Almer	4ee38ed7f9	x86/float_dsp: zero extend len from ff_butterflies_float_sse implicitly Reviewed-by: Christophe Gisquet <christophe.gisquet@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	2016-01-08 16:14:27 -03:00
James Almer	7f520524f6	x86/float_dsp: remove len check from ff_butterflies_float_sse The function documentation explicitly mentions it needs to be a multiple of 4. Reviewed-by: Christophe Gisquet <christophe.gisquet@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	2016-01-08 16:14:24 -03:00
James Almer	f4c1a48483	x86/intmath: add sse optimized av_clipf and av_clipd Reviewed-by: Michael Niedermayer <michael@niedermayer.cc> Reviewed-by: Ronald S. Bultje <rsbultje@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	2016-01-07 14:24:01 -03:00
Matt Oliver	e9ec28c95e	avutil/x86/bswap: Remove warning about bswap intrinsics with msvc. Signed-off-by: Matt Oliver <protogonoi@gmail.com>	2015-11-23 23:03:32 +11:00
Matt Oliver	58d32c00be	avutil/x86/intmath: Fix intrinsic header include when using newer gcc with older icc. Signed-off-by: Matt Oliver <protogonoi@gmail.com>	2015-11-12 16:54:08 +11:00
Matt Oliver	9b9c9ef3b4	avutil/x86/bswap: Add msvc bswap instrinsics. This adds msvc optimisations as well as fixing an error in icl whereby it will generate invalid code otherwise. Signed-off-by: Matt Oliver <protogonoi@gmail.com>	2015-11-12 16:53:44 +11:00
Matt Oliver	9105399060	avutil/x86/intmath: Disable use of tzcnt on older intel compilers. ICC versions older than atleast 12.1.6 dont have the tzcnt intrinsics. Reviewed-by: Michael Niedermayer <michael@niedermayer.cc> Signed-off-by: Matt Oliver <protogonoi@gmail.com>	2015-11-11 10:18:08 +11:00
Matt Oliver	f984174512	avutil/x86/intmath: Correct intrinsic headers for older compilers. Signed-off-by: Matt Oliver <protogonoi@gmail.com>	2015-11-09 21:40:33 +11:00
Matt Oliver	bff009697d	avutil/x86/intmath: Add missing header. Signed-off-by: Matt Oliver <protogonoi@gmail.com>	2015-11-01 02:11:29 +11:00
Matt Oliver	6c6ac9cb17	avutil/x86/intmath: Use tzcnt in place of bsf. Signed-off-by: Matt Oliver <protogonoi@gmail.com>	2015-10-31 23:11:32 +11:00
Rodger Combs	1e477a970f	lavu: add AESNI CPU flag	2015-10-28 04:23:14 -05:00
Matt Oliver	b0bb1dc62d	lavu/intmath.h: Move x86 only msvc/icl functions to x86 specific header. Signed-off-by: Matt Oliver <protogonoi@gmail.com>	2015-10-19 13:40:51 +11:00
Matt Oliver	216cc1f6fe	lavu/intmath.h: Add msvc/icl ctzll optimisations. Signed-off-by: Matt Oliver <protogonoi@gmail.com>	2015-10-19 13:40:27 +11:00
Henrik Gramner	17710550c4	x86inc: Make cpuflag() and notcpuflag() return 0 or 1 Makes it possible to use them in arithmetic expressions.	2015-10-01 18:14:12 +02:00
James Almer	36e1665d3d	avutil/attributes: add AV_GCC_VERSION_AT_MOST Reviewed-by: Michael Niedermayer <michaelni@gmx.at> Signed-off-by: James Almer <jamrial@gmail.com>	2015-09-18 12:41:29 -03:00
James Almer	d5f8a642f6	x86: port PSIGNW to cpuflags Reviewed-by: Ronald S. Bultje <rsbultje@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	2015-09-11 23:27:03 -03:00
Ganesh Ajjanagadde	531b0a316b	avutil/x86/asm: rename REG_SP to REG_sp REG_SP is defined by Solaris system headers. This fixes a sea of warnings while building on Solaris: http://fate.ffmpeg.org/report.cgi?time=20150820233505&slot=x86-opensolaris-gcc4.3 Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com> Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>	2015-08-22 02:56:53 +02:00
Henrik Gramner	44b4444120	x86inc: Various minor backports from x264 Signed-off-by: Anton Khirnov <anton@khirnov.net>	2015-08-13 07:46:24 +02:00
Henrik Gramner	ab43beefab	x86inc: Drop SECTION_TEXT macro The .text section is already 16-byte aligned by default on all supported platforms so `SECTION_TEXT` isn't any different from `SECTION .text`. Signed-off-by: Anton Khirnov <anton@khirnov.net>	2015-08-11 11:12:01 +02:00
Henrik Gramner	1c6bb81328	x86inc: Disable vpbroadcastq workaround in newer yasm versions The bug was fixed in 1.3.0, so only perform the workaround in earlier versions. Signed-off-by: Anton Khirnov <anton@khirnov.net>	2015-08-11 11:11:27 +02:00
Christophe Gisquet	f5e486f6f8	x86inc: Fix instantiation of YMM registers Signed-off-by: Henrik Gramner <henrik@gramner.com> Signed-off-by: Anton Khirnov <anton@khirnov.net>	2015-08-11 11:09:08 +02:00
Anton Mitrofanov	b114d28a18	x86inc: warn when instructions incompatible with current cpuflags are used Signed-off-by: Henrik Gramner <henrik@gramner.com> Signed-off-by: Anton Khirnov <anton@khirnov.net>	2015-08-11 11:07:18 +02:00
Henrik Gramner	9f1245eb96	x86inc: Support arbitrary stack alignments Change ALLOC_STACK to always align the stack before allocating stack space for consistency. Previously alignment would occur either before or after allocating stack space depending on whether manual alignment was required or not. Signed-off-by: Anton Khirnov <anton@khirnov.net>	2015-08-11 11:04:11 +02:00
Anton Mitrofanov	8c75ba55a4	x86inc: warn if XOP integer FMA instruction emulation is impossible Emulation requires a temporary register if arguments 1 and 4 are the same; this doesn't obey the semantics of the original instruction, so we can't emulate that in x86inc. Also add pmacsdql emulation. Signed-off-by: Henrik Gramner <henrik@gramner.com> Signed-off-by: Anton Khirnov <anton@khirnov.net>	2015-08-11 11:02:27 +02:00
Anton Mitrofanov	8db0f71b49	x86inc: warn if XOP integer FMA instruction emulation is impossible Signed-off-by: Henrik Gramner <henrik@gramner.com>	2015-08-05 16:15:40 +02:00
Henrik Gramner	f0b7882ceb	x86inc: Drop SECTION_TEXT macro The .text section is already 16-byte aligned by default on all supported platforms so `SECTION_TEXT` isn't any different from `SECTION .text`.	2015-08-04 20:13:09 +02:00
Henrik Gramner	826790f596	x86inc: Support arbitrary stack alignments Change ALLOC_STACK to always align the stack before allocating stack space for consistency. Previously alignment would occur either before or after allocating stack space depending on whether manual alignment was required or not.	2015-08-04 20:13:09 +02:00
James Almer	5750d6c5e9	x86: move XOP emulation code back to x86inc Only two functions that use xop multiply-accumulate instructions where the first operand is the same as the fourth actually took advantage of the macros. This further reduces differences with x264's x86inc. Reviewed-by: Ronald S. Bultje <rsbultje@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	2015-08-03 17:11:13 -03:00
Henrik Gramner	127203ba5a	x86inc: Various minor backports from x264 Reviewed-by: "Ronald S. Bultje" <rsbultje@gmail.com> Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>	2015-08-03 04:08:33 +02:00
Henrik Gramner	f151fbd9e5	x86inc: Disable vpbroadcastq workaround in newer yasm versions The bug was fixed in 1.3.0, so only perform the workaround in earlier versions. Reviewed-by: "Ronald S. Bultje" <rsbultje@gmail.com> Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>	2015-08-03 03:13:20 +02:00
James Almer	4d2c014a8f	x86/float_dsp: add missing colon to labels Silences warnings with Nasm Signed-off-by: James Almer <jamrial@gmail.com>	2015-07-26 02:51:08 -03:00
James Almer	bd48764532	avutil/x86/bswap: force inline asm versions with ICC Recent ICC versions that define GCC as >= 4.5 (like ICC 13) apparently can't optimize the generic C versions of av_bswap*() on their own. Reviewed-by: Michael Niedermayer <michaelni@gmx.at> Signed-off-by: James Almer <jamrial@gmail.com>	2015-07-18 20:48:09 -03:00
Michael Niedermayer	2ecbf44f21	Merge commit 'd1a6cb195f610978ba5d2351e60f938f7f261d59' * commit 'd1a6cb195f610978ba5d2351e60f938f7f261d59': x86: Serialize rdtsc in read_time() Merged-by: Michael Niedermayer <michaelni@gmx.at>	2015-07-09 12:28:09 +02:00
Henrik Gramner	d1a6cb195f	x86: Serialize rdtsc in read_time() Improves the accuracy of measurements, especially in short sections. To quote the Intel 64 and IA-32 Architectures Software Developer's Manual: "The RDTSC instruction is not a serializing instruction. It does not necessarily wait until all previous instructions have been executed before reading the counter. Similarly, subsequent instructions may begin execution before the read operation is performed. If software requires RDTSC to be executed only after all previous instructions have completed locally, it can either use RDTSCP (if the processor supports that instruction) or execute the sequence LFENCE;RDTSC." SSE2 is a requirement for lfence so only use it on SSE2-capable systems. Prefer lfence;rdtsc over rdtscp since rdtscp is supported on fewer systems. Signed-off-by: Luca Barbato <lu_zero@gentoo.org>	2015-07-09 00:10:13 +02:00
James Almer	93e7b7fb34	avutil/x86/intmath: add missing check for inline assembly Signed-off-by: James Almer <jamrial@gmail.com>	2015-06-27 14:33:53 -03:00
James Almer	1e51e517be	avutil/x86/intmath: use bzhi gcc builtin in av_mod_uintp2() Signed-off-by: James Almer <jamrial@gmail.com>	2015-06-27 12:56:55 -03:00
James Almer	c16e99e3b3	x86: check for AV_CPU_FLAG_AVXSLOW where useful Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2015-06-01 00:15:35 +02:00
Michael Niedermayer	16c430e8ef	Merge commit 'cae39851201b7781f1262e1c23627b45e6e80bb4' * commit 'cae39851201b7781f1262e1c23627b45e6e80bb4': x86: Add helper macros to check for slow cpuflags Merged-by: Michael Niedermayer <michaelni@gmx.at>	2015-05-31 23:59:48 +02:00
James Almer	cae3985120	x86: Add helper macros to check for slow cpuflags Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Luca Barbato <lu_zero@gentoo.org>	2015-05-31 12:07:11 +02:00
James Almer	d68c05380c	x86: check for AV_CPU_FLAG_AVXSLOW where useful Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Luca Barbato <lu_zero@gentoo.org>	2015-05-31 12:07:11 +02:00
James Almer	f7cafb5d02	x86: add AV_CPU_FLAG_AVXSLOW flag Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Luca Barbato <lu_zero@gentoo.org>	2015-05-31 12:07:11 +02:00
Timothy Gu	dd4d709be7	x86inc: Clear __SECT__ Silences warning(s) like: libavcodec/x86/fft.asm:93: warning: section flags ignored on section redeclaration The cause of this warning is that because `struc` and `endstruc` attempts to revert to the previous section state [1]. The section state is stored in the macro __SECT__, defined by x86inc.asm to be `.note.GNU-stack ...`, through the `SECTION` directive [2]. Thus, the `.note.GNU-stack` section is defined twice (once in x86inc.asm, once during `endstruc`), causing the warning. That is the first part of the commit: using the primitive `[section]` format for .note.GNU-stack etc., which does not update `__SECT__` [2]. That fixes only half of the problem. Even without any `SECTION` directives, `__SECT__` is predefined as `.text`, which conflicting with the later `SECTION_TEXT` (which expands to `.text align=16`). [1]: http://www.nasm.us/doc/nasmdoc6.html#section-6.4 [2]: http://www.nasm.us/doc/nasmdoc6.html#section-6.3 Signed-off-by: Luca Barbato <lu_zero@gentoo.org>	2015-05-28 11:40:15 +02:00
Timothy Gu	204b228a1d	x86inc: Clear __SECT__ This commit silences warning(s) like: libavcodec/x86/fft.asm:93: warning: section flags ignored on section redeclaration The cause of this warning is that because `struc` and `endstruc` attempts to revert to the previous section state [1]. The section state is stored in the macro __SECT__, defined by x86inc.asm to be `.note.GNU-stack ...`, through the `SECTION` directive [2]. Thus, the `.note.GNU-stack` section is defined twice (once in x86inc.asm, once during `endstruc`), causing the warning. That is the first part of the commit: using the primitive `[section]` format for .note.GNU-stack etc., which does not update `__SECT__` [2]. That fixes only half of the problem. Even without any `SECTION` directives, `__SECT__` is predefined as `.text`, which conflicting with the later `SECTION_TEXT` (which expands to `.text align=16`). [1]: http://www.nasm.us/doc/nasmdoc6.html#section-6.4 [2]: http://www.nasm.us/doc/nasmdoc6.html#section-6.3 Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2015-05-28 00:08:37 +02:00
James Almer	c312bfac4c	x86/cpu: add AV_CPU_FLAG_AVXSLOW flag Reviewed-by: Michael Niedermayer <michaelni@gmx.at> Signed-off-by: James Almer <jamrial@gmail.com>	2015-05-27 03:31:11 -03:00
Michael Niedermayer	d630f38f47	avutil/x86/Makefile: fix conditional x86/emms.o build Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2015-04-09 01:12:51 +02:00
Ronald S. Bultje	b926f02e81	avutil/x86/Makefile: Make building and linking of emms.c conditional Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2015-04-08 17:25:35 +02:00
James Almer	60b9373dbd	libavutil: add bmi2 optimized av_mod_uintp2 Reviewed-by: Michael Niedermayer <michaelni@gmx.at> Signed-off-by: James Almer <jamrial@gmail.com>	2015-03-20 15:47:43 -03:00
Peter Cordes	9e5687adf2	pixelutils: Comment on (lack of) sad_8x8_sse2 Signed-off-by: Peter Cordes <peter@cordes.ca>	2015-03-04 21:58:53 +01:00
James Almer	bc65abc8d7	libavutil: add x86 optimized av_popcount Reviewed-by: Ronald S. Bultje <rsbultje@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	2015-02-25 19:58:00 -03:00
Christophe Gisquet	d9293c776e	x86inc: Correctly warn on use of SSE2 instructions in SSE functions SSE2 instructions that are XMM-implementations of pre-existing MMX/MMX2 instructions did not issue warnings when used in SSE functions. Handle it by also checking the register type when such instructions are used. Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2015-02-17 12:35:58 +01:00
Christophe Gisquet	e93d3a22cb	x86: lavu/x264asm: fix ymm register instantiation This mimicks what is done for the other instruction sets. Tested-by: James Almer <jamrial@gmail.com> Tested-by: Mickaël Raulet <mraulet@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2015-02-04 00:18:29 +01:00
James Darnley	12120174ce	lavu/x86/x86inc: deprecate INIT_AVX The same can be done with INIT_XMM avx Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2015-02-02 01:09:16 +01:00
Anton Mitrofanov	a1684311b3	x264asm: warn when inappropriate instruction used in function with specified cpuflags Requested-by: Christophe Gisquet <christophe.gisquet@gmail.com> Requested-by: "Ronald S. Bultje" <rsbultje@gmail.com>	2015-02-02 00:06:14 +01:00
James Almer	37b35feb64	x86/swr: add SSE2/AVX pack_8ch functions Reviewed-by: Michael Niedermayer <michaelni@gmx.at> Reviewed-by: Ronald S. Bultje <rsbultje@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	2014-12-30 23:05:27 -03:00
Kieran Kunhya	9a738c27dc	v210enc: Add SIMD optimised 8-bit and 10-bit encoders Signed-off-by: Michael Niedermayer <michaelni@gmx.at> Signed-off-by: Vittorio Giovara <vittorio.giovara@gmail.com>	2014-12-05 13:03:49 +00:00
Kieran Kunhya	36091742d1	v210enc: Add SIMD optimised 8-bit and 10-bit encoders Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-11-26 20:30:47 +01:00
Michael Niedermayer	579a0fdc21	avutil/lls: Make unchanged function arguments const Reviewed-by: Paul B Mahol <onemda@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-09-28 19:32:07 +02:00
lvqcl	e58fc44649	avutil/x86/cpu: fix cpuid sub-leaf selection Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-09-27 13:21:31 +02:00
Henrik Gramner	f629705b02	x86inc: Make INIT_CPUFLAGS support an arbitrary number of cpuflags Previously there was a limit of two cpuflags. Signed-off-by: Diego Biurrun <diego@biurrun.de>	2014-09-09 02:00:25 -07:00
Loren Merritt	ec217218c2	x86inc: Free up variable name "n" in global namespace Signed-off-by: Diego Biurrun <diego@biurrun.de>	2014-09-09 02:00:19 -07:00
Henrik Gramner	176a0fca3f	x86inc: Make ym# behave the same way as xm# This makes more sense for future implementations of templates with zmm registers. Signed-off-by: Diego Biurrun <diego@biurrun.de>	2014-09-09 01:45:14 -07:00
Henrik Gramner	428aa14a48	x86inc: Make INIT_CPUFLAGS support an arbitrary number of cpuflags Previously there was a limit of two cpuflags. Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-09-05 14:06:03 +02:00
Henrik Gramner	720c21d11f	x86inc: Make ym# behave the same way as xm# This makes more sense for future implementations of templates with zmm registers. Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-09-05 01:55:28 +02:00
Loren Merritt	a4dbabc8b3	x86inc: free up variable name "n" in global namespace Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-09-05 01:41:50 +02:00
Clément Bœsch	554d819062	avutil/pixelutils: faster pixelutils_sad_16x16 501 to 439 decicycles. See `45c7f3997e`.	2014-08-23 20:12:56 +02:00
Clément Bœsch	45c7f3997e	avutil/pixelutils: faster pixelutils_sad_[au]_16x16 ~560 → ~500 decicycles This is following the comments from Michael in https://ffmpeg.org/pipermail/ffmpeg-devel/2014-August/160599.html Using 2 registers for accumulator didn't help. On the other hand, some re-ordering between the movs and psadbw allowed going ~538 to ~500.	2014-08-23 10:18:53 +02:00
Michael Niedermayer	70b8668fb5	drop LLS1, rename LLS2 to LLS Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-08-09 23:20:31 +02:00
Clément Bœsch	28a2107a8d	avutil: add pixelutils API	2014-08-05 21:05:52 +02:00
James Almer	d0f56ca071	x86/hevc_deblock: improve 8bit transpose store macros Up to four instructions less depending on function and instruction set. Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-08-03 04:24:15 +02:00
James Almer	1ace9573dc	x86/hevc_idct: replace old and unused idct functions Only 8-bit and 10-bit idct_dc() functions are included (adding others should be trivial). Benchmarks on an Intel Core i5-4200U: idct8x8_dc SSE2 MMXEXT C cycles 22 26 57 idct16x16_dc AVX2 SSE2 C cycles 27 32 249 idct32x32_dc AVX2 SSE2 C cycles 62 126 1375 Signed-off-by: James Almer <jamrial@gmail.com> Reviewed-by: Mickaël Raulet <mraulet@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-07-26 18:00:11 +02:00
Michael Niedermayer	8d0c7031a8	Merge commit '79793f833784121d574454af4871866576c0749d' * commit '79793f833784121d574454af4871866576c0749d': Update Fiona's name in copyright statements. Merged-by: Michael Niedermayer <michaelni@gmx.at>	2014-07-01 15:43:40 +02:00
Diego Biurrun	79793f8337	Update Fiona's name in copyright statements.	2014-07-01 03:26:51 -07:00
Christophe Gisquet	9107612818	x86util: add and use RSHIFT/LSHIFT macros Those macros take a byte number as shift argument, as this argument differs between MMX and SSE2 instructions. Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-06-15 13:19:27 +02:00
James Almer	85065d2a7c	x86/float_dsp: add missing femms It was lost during the port. Should fix fate on 3dnowext machines. Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-06-08 20:06:28 +02:00
James Almer	dcaf9660b6	x86/float_dsp: port vector_fmul_window to yasm Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-06-08 12:41:32 +02:00
James Almer	fc8db12a73	x86/vp9: inital AVX2 intra_pred tos3k-vp9-b10000.webm on a Core i5-4200U @1.6GHz 1219 decicycles in ff_vp9_ipred_dc_32x32_ssse3, 131070 runs, 2 skips 439 decicycles in ff_vp9_ipred_dc_32x32_avx2, 131070 runs, 2 skips 3570 decicycles in ff_vp9_ipred_dc_top_32x32_ssse3, 4096 runs, 0 skips 2494 decicycles in ff_vp9_ipred_dc_top_32x32_avx2, 4096 runs, 0 skips 1419 decicycles in ff_vp9_ipred_dc_left_32x32_ssse3, 16384 runs, 0 skips 717 decicycles in ff_vp9_ipred_dc_left_32x32_avx2, 16384 runs, 0 skips 2737 decicycles in ff_vp9_ipred_tm_32x32_avx, 1024 runs, 0 skips 2088 decicycles in ff_vp9_ipred_tm_32x32_avx2, 1024 runs, 0 skips 3090 decicycles in ff_vp9_ipred_v_32x32_avx, 512 runs, 0 skips 2226 decicycles in ff_vp9_ipred_v_32x32_avx2, 512 runs, 0 skips 1565 decicycles in ff_vp9_ipred_h_32x32_avx, 1024 runs, 0 skips 922 decicycles in ff_vp9_ipred_h_32x32_avx2, 1024 runs, 0 skips Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-06-08 02:37:20 +02:00
Christophe Gisquet	2267003981	x86: hpeldsp: better factorization Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-05-29 21:47:40 +02:00
James Almer	561bfc85eb	x86/dsputilenc: implement SSE2 versions of pix_{sum16, norm1} Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-05-28 23:29:34 +02:00
Matt Oliver	1898c2f49d	inline asm: fix arrays as named constraints. Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-05-07 15:02:45 +02:00
James Almer	3b06208a57	x86/float_dsp: remove duplicated code from vector_dmul_scalar Use the xm# and ym# aliases as they remain in sync with m# after a SWAP. No actual changes to the assembly. Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-04-19 14:21:51 +02:00
James Almer	76ed71a72b	x86: move horizontal add macros to x86util Also port relevant AVX2/XOP optimizations from x264 with permission to relicense to LGPL from the corresponding authors Signed-off-by: James Almer <jamrial@gmail.com> Reviewed-by: "Ronald S. Bultje" <rsbultje@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-04-17 14:15:09 +02:00
James Almer	11b36b1ee0	x86/float_dsp: unroll loop in vector_fmac_scalar ~6% faster SSE2 performance. AVX/FMA3 are unaffected. Signed-off-by: James Almer <jamrial@gmail.com> Reviewed-by: Christophe Gisquet <christophe.gisquet@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-04-16 18:36:52 +02:00
James Almer	3b808900af	x86/float_dsp: use SWAP in vector_fmac_scalar Win64 The mova is unnecessary Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-04-16 15:46:21 +02:00
James Almer	2d9821a208	x86/cpu: check for OS support before enabling AVX2 AV_CPU_FLAG_AVX is enabled at this point only if there's OS support. Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-03-25 17:56:43 +01:00
Matt Oliver	8236747511	Automatically change MANGLE() into named inline asm operands when direct symbol reference in inline asm are not supported. This is part of the patch-set for intel C inline asm on windows support Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-03-18 23:39:30 +01:00
James Almer	7d7487e85c	x86/float_dsp: add ff_vector_{fmul_add, fmac_scalar}_fma3 ~7% faster than AVX Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-03-13 04:34:05 +01:00
Michael Niedermayer	4159f702a7	avutil/timer: Fix units for x86 after `c708b54033` Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-03-09 15:22:02 +01:00
James Almer	3f3d748cab	x86: Move XOP emulation to x86util We need the emulation to support the cases where the first argument is the same as the fourth. To achieve this a fifth argument working as a temporary may be needed. Emulation that doesn't obey the original instruction semantics can't be in x86inc. Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-24 08:30:19 +01:00
Michael Niedermayer	bd8d73ea8b	Merge remote-tracking branch 'qatar/master' * qatar/master: x86: add detection for Bit Manipulation Instruction sets Conflicts: libavutil/x86/cpu.c See: `0bc3de19ff` Merged-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-23 22:52:58 +01:00

1 2 3 4 5 ...

564 Commits