FFmpeg

Commit Graph

Author	SHA1	Message	Date
Lynne	bbe95f7353	x86: replace explicit REP_RETs with RETs From x86inc: > On AMD cpus <=K10, an ordinary ret is slow if it immediately follows either > a branch or a branch target. So switch to a 2-byte form of ret in that case. > We can automatically detect "follows a branch", but not a branch target. > (SSSE3 is a sufficient condition to know that your cpu doesn't have this problem.) x86inc can automatically determine whether to use REP_RET rather than REP in most of these cases, so impact is minimal. Additionally, a few REP_RETs were used unnecessary, despite the return being nowhere near a branch. The only CPUs affected were AMD K10s, made between 2007 and 2011, 16 years ago and 12 years ago, respectively. In the future, everyone involved with x86inc should consider dropping REP_RETs altogether.	2 years ago
Lynne	90c17a05aa	x86/tx_float: fix stray change in 15xM FFT and replace imul->lea Thanks to rorgoroth for bisecting and kurosu for the lea suggestion.	2 years ago
Lynne	87bae6b018	lavu/tx: refactor to explicitly track and convert lookup table order Necessary for generalizing PFAs.	2 years ago
Lynne	fab97faf02	x86/tx_float: implement striding in fft_15xM	2 years ago
Lynne	92100eee5b	x86/tx_float_init: properly specify the supported factors of 15xM FFTs Only powers of two are currently supported.	2 years ago
Lynne	cc1df4045e	x86/tx_float: add a standalone 15-point AVX2 transform Enables its use everywhere else in the framework.	2 years ago
Lynne	877e575b5d	x86/tx_float: optimize and macro out FFT15	2 years ago
Johannes Kauffmann	a11e745b97	lavu/fixed_dsp: add missing av_restrict qualifiers The butterflies_fixed function pointer declaration specifies av_restrict for the first two pointer arguments. So the corresponding function definitions should honor this declaration. MSVC emits warning C4113 for this. Signed-off-by: Anton Khirnov <anton@khirnov.net>	2 years ago
Lynne	f21899db7d	x86/tx_float: enable AVX-only split-radix FFT codelets Sandy Bridge, Ivy Bridge and Bulldozer cores don't support FMA3.	2 years ago
James Almer	d2f482965f	x86/tx_float: fix some symbol names Should fix compilation on MacOS Signed-off-by: James Almer <jamrial@gmail.com>	2 years ago
James Almer	0d8f43c74d	x86/tx_float: change a condition in a preprocessor check Fixes compilation with yasm. Signed-off-by: James Almer <jamrial@gmail.com>	2 years ago
James Almer	750f378bec	x86/tx_float: add missing preprocessor wrapper for AVX2 functions Fixes compilation with old assemblers. Signed-off-by: James Almer <jamrial@gmail.com>	2 years ago
Lynne	74e8541bab	x86/tx_float: generalize iMDCT To support non-aligned buffers during the post-transform step, just iterate backwards over the array. This allows using the 15xN-point FFT, with which the speed is 2.1 times faster than our old libavcodec implementation.	2 years ago
Lynne	ace42cf581	x86/tx_float: add 15xN PFA FFT AVX SIMD ~4x faster than the C version. The shuffles in the 15pt dim1 are seriously expensive. Not happy with it, but I'm contempt. Can be easily converted to pure AVX by removing all vpermpd/vpermps instructions.	2 years ago
Lynne	3241e9225c	x86/tx_float: adjust internal ASM call ABI again There are many ways to go about it, and this one seems optimal for both MDCTs and PFA FFTs without requiring excessive instructions or stack usage.	2 years ago
Lynne	4ba68639ca	x86/tx_float: add asm call versions of the 2pt and 4pt transforms Verified to be working.	2 years ago
Lynne	892548e6a1	x86/tx_float: fully support 128bit regs in LOAD64_LUT The gather path didn't support 128bit registers. It's not faster on Zen 3, but it's here for completeness.	2 years ago
Lynne	af42bb3d61	x86/tx_float: simplify and describe the intra-asm call convention	2 years ago
James Almer	bda3a9faf4	x86/float_dsp: use three operand form for some instructions Fixes compilation with old yasm Signed-off-by: James Almer <jamrial@gmail.com>	2 years ago
Paul B Mahol	72acff9f59	avutil/x86/float_dsp: add fma3 for scalarproduct	2 years ago
Andreas Rheinhardt	29c4c0886d	avutil/x86/intreadwrite: Add ability to detect whether MMX code is used It can be used to call emms_c() only when needed. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@outlook.com>	2 years ago
James Almer	f4097e4c1f	x86/tx_float: add missing check for AVX2 Fixes compilation with old yasm. Signed-off-by: James Almer <jamrial@gmail.com>	2 years ago
James Almer	74f5fb6db8	x86/tx_float: set all operands for shufps Fixes compilation with AVX2 enabled yasm. Signed-off-by: James Almer <jamrial@gmail.com>	2 years ago
Martin Storsjö	e4759fa951	x86/tx_float: Fix building for platforms with a symbol prefix This fixes building for x86 macOS (both i386 and x86_64) and i386 windows. Signed-off-by: Martin Storsjö <martin@martin.st>	2 years ago
Lynne	4537d9554d	x86/tx_float: implement inverse MDCT AVX2 assembly This commit implements an iMDCT in pure assembly. This is capable of processing any mod-8 transforms, rather than just power of two, but since power of two is all we have assembly for currently, that's what's supported. It would really benefit if we could somehow use the C code to decide which function to jump into, but exposing function labels from assebly into C is anything but easy. The post-transform loop could probably be improved. This was somewhat annoying to write, as we must support arbitrary strides during runtime. There's a fast branch for stride == 4 bytes and a slower one which uses vgatherdps. Zen 3 benchmarks for stride == 4 for old (av_imdct_half) vs new (av_tx): 128pt: 2811 decicycles in av_tx (imdct),16775916 runs, 1300 skips 3082 decicycles in av_imdct_half,16776751 runs, 465 skips 256pt: 4920 decicycles in av_tx (imdct),16775820 runs, 1396 skips 5378 decicycles in av_imdct_half,16776411 runs, 805 skips 512pt: 9668 decicycles in av_tx (imdct),16775774 runs, 1442 skips 10626 decicycles in av_imdct_half,16775647 runs, 1569 skips 1024pt: 19812 decicycles in av_tx (imdct),16777144 runs, 72 skips 23036 decicycles in av_imdct_half,16777167 runs, 49 skips	2 years ago
Lynne	2425d5cd7e	x86/tx_float: add support for calling assembly functions from assembly Needed for the next patch. We get this for the extremely small cost of a branch on _ns functions, which wouldn't be used anyway with assembly.	2 years ago
Lynne	98b32ef462	x86/tx_float: save a branch during coefficient deinterleaving Directly branch into the special 64-point deinterleave subroutine rather than going through the general deinterleave. 64-point transform timings on Zen 3: Before: 1974 decicycles in av_tx (fft),16776864 runs, 352 skips After: 1956 decicycles in av_tx (fft),16775378 runs, 1838 skips	2 years ago
Andreas Rheinhardt	2718a3be1f	avutil/x86/float_dsp: Remove obsolete 3dnowext function x64 always has MMX, MMXEXT, SSE and SSE2 and this means that some functions for MMX, MMXEXT, SSE and 3dnow are always overridden by other functions (unless one e.g. explicitly disables SSE2). So given that the only systems which benefit from ff_vector_fmul_window_3dnowext are truely ancient 32bit AMD x86s it is removed. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@outlook.com>	2 years ago
Andreas Rheinhardt	ea043cc53e	avutil/x86/pixelutils: Remove obsolete MMX(EXT) functions x64 always has MMX, MMXEXT, SSE and SSE2 and this means that some functions for MMX, MMXEXT, SSE and 3dnow are always overridden by other functions (unless one e.g. explicitly disables SSE2). So given that the only systems which benefit from the 8x8 MMX (overridden by MMXEXT) or the 16x16 MMXEXT (overridden by SSE2) are truely ancient 32bit x86s they are removed. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@outlook.com>	2 years ago
Lynne	27cffd16aa	x86/tx_float: replace fft_sr_avx with fft_sr_fma3 When the SLOW_GATHER flag was added to the AVX2 version, this made FMA3-features not enabled on Zen CPUs. As FMA3 adds 6-7% across all platforms that support it, in the interest of saving space, this commit removes the AVX version and replaces it with an FMA3 version. The only CPUs affected are Sandy Bridge and Bulldozer, which have AVX support, but no FMA3 support. In the future, if there's a demand for it, a version of the function duplicated for AVX can be added.	3 years ago
Lynne	0938ff9701	x86/tx_float: improve temporary register allocation for loads On Zen 3: Before: 1484285 decicycles in av_tx (fft), 131072 runs, 0 skips After: 1415243 decicycles in av_tx (fft), 131072 runs, 0 skips	3 years ago
Lynne	19c0bb2aa9	x86/tx_float: add AV_CPU_FLAG_AVXSLOW/SLOW_GATHER flags where appropriate	3 years ago
Lynne	9e94c35941	Revert "x86/tx_float: remove vgatherdpd usage" This reverts commit `82a68a8771`. Smarter slow ISA penalties makes gathers still useful. The intention is to use gathers with the final stage of non-ptwo iMDCTs, where they give benefit.	3 years ago
Lynne	82a68a8771	x86/tx_float: remove vgatherdpd usage Its performance loss ranges from either being just as fast as individual loads (Skylake), a few percent slower (Alderlake), 8% slower (Zen 3), to completely disasterous (older/other CPUs). Sadly, gathers never panned out fast on x86, even with the benefit of time and implementation experience. This also saves a register, as there's no need to fill out an additional register mask. Zen 3 (16384-point transform): Before: 1561050 decicycles in av_tx (fft), 131072 runs, 0 skips After: 1449621 decicycles in av_tx (fft), 131072 runs, 0 skips Alderlake: 2% slower on big transforms (65536), to 1% (131072), to a few percent for smaller sizes.	3 years ago
Wu Jianhua	f629ea2e18	avutil/cpu: add AVX512 Icelake flag Signed-off-by: Wu Jianhua <jianhua.wu@intel.com> Reviewed-by: Henrik Gramner <henrik@gramner.com> Signed-off-by: James Almer <jamrial@gmail.com>	3 years ago
Andreas Rheinhardt	636631d9db	Remove unnecessary libavutil/(avutil\|common\|internal).h inclusions Some of these were made possible by moving several common macros to libavutil/macros.h. While just at it, also improve the other headers a bit. Reviewed-by: Martin Storsjö <martin@martin.st> Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@outlook.com>	3 years ago
Andreas Rheinhardt	6c694074e1	avutil/x86/emms: Don't unnecessarily include lavu/cpu.h Only include it if it is needed, namely if __MMX__ is undefined. X86 is currently the only arch where lavu/cpu.h is basically automatically included (for internal development): #if ARCH_X86 is true, lavu/internal.h (which is basically included everywhere) includes lavu/x86/emms.h which can mask missing inclusions of lavu/cpu.h if the developer works on x86/x64. This has happened in `8e825ec3ab` and also earlier (see `6d2365882f`). By including said header only if necessary ordinary developer machines will behave like non-x86 arches, so that missing inclusions of cpu.h won't go unnoticed any more. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@outlook.com>	3 years ago
Alexander Kanavin	91326dc942	libavutil: include assembly with full path from source root Otherwise nasm writes the full host-specific paths into .o output, which breaks binary reproducibility. Signed-off-by: Alexander Kanavin <alex.kanavin@gmail.com> Signed-off-by: Anton Khirnov <anton@khirnov.net>	3 years ago
Lynne	3bbe9c5e38	lavu/tx: refactor assembly codelet definition This commit does some refactoring to make defining assembly codelets smaller, and fixes compiler redefinition warnings. It also allows for other assembly versions to reuse the same boilerplate code as x86. Finally, it also adds the out_of_place flag to all assembly codelets. This changes nothing, as out-of-place operation was assumed to be available anyway, but this makes it more explicit.	3 years ago
Lynne	2e82c61055	x86/tx_float: avoid redefining macros FFT16_FN was used for fft8 and for fft16 afterwards.	3 years ago
Lynne	35080149ef	x86/tx_float: mark AVX2 functions as AVXSLOW Makes Bulldozer prefer AVX functions rather than AVX2, which are 64% slower: AVX: 117653 decicycles in av_tx (fft), 1048535 runs, 41 skips AVX2: 193385 decicycles in av_tx (fft), 1048561 runs, 15 skips The only difference between both is that vgatherdpd is used in the former. We don't want to mark them with the new SLOW_GATHER flag however, since gathers are still faster on Haswell/Zen 2/3 than plain loads.	3 years ago
Lynne	6c397f6bb5	x86/tx_float: add missing FF_TX_OUT_OF_PLACE flag to functions This caused smaller length dedicated transforms to not be picked up.	3 years ago
Lynne	9787005846	x86/tx_float: do not build tx_float_init.c if x86 assembly is disabled This broke builds with --disable-mmx, which also disabled assembly entirely, but ARCH_X86 was still true, so the init file tried to find assembly that didn't exist. Instead of checking for architecture, check if external x86 assembly is enabled.	3 years ago
Lynne	28bff6ae54	x86/tx_float: add permute-free FFT versions These are used in the PFA transforms and MDCTs.	3 years ago
Lynne	ef4bd81615	lavu/tx: rewrite internal code as a tree-based codelet constructor This commit rewrites the internal transform code into a constructor that stitches transforms (codelets). This allows for transforms to reuse arbitrary parts of other transforms, and allows transforms to be stacked onto one another (such as a full iMDCT using a half-iMDCT which in turn uses an FFT). It also permits for each step to be individually replaced by assembly or a custom implementation (such as an ASIC).	3 years ago
James Almer	8c2d2fd6cc	avutil/cpu: move slow gather checks below in the function Put them together with other similar slow flag checks. Signed-off-by: James Almer <jamrial@gmail.com>	3 years ago
Alan Kelly	ffbab99f2c	libavutil/cpu: Add AV_CPU_FLAG_SLOW_GATHER. This flag is set on Haswell and earlier and all AMD cpus.	3 years ago
James Almer	67b92d68c6	x86/intmath: add VEX encoded versions of av_clipf() and av_clipd() Prevents mixing inlined SSE instructions and AVX instructions when the compiler generates the latter. Signed-off-by: James Almer <jamrial@gmail.com>	3 years ago
Mark Reid	c3502f4f75	libavutil/common: clip nan value to amin Changes av_clipf to return amin if a is nan. Before if a is nan av_clipf_c returned nan and av_clipf_sse would return amax. Now the both should behave the same. This works because nan > amin is false. The max(nan, amin) will be amin. Signed-off-by: James Almer <jamrial@gmail.com>	3 years ago
Lynne	997f9bdb99	x86/tx_float: correctly load the transform length The field is a standard field, yet we were loading it as if it was a quadword. This worked for forward transforms by chance, but broke when the transform was inverse. checkasm couldn't catch that because we only test forward transforms, which are identical to inverse transforms but with a different revtab.	3 years ago

1 2 3 4 5 ...

565 Commits (2ad468ed1f7a7c67f1a2ab37c9e304b0011f0aae)