FFmpeg

Commit Graph

Author	SHA1	Message	Date
Anton Khirnov	580e168a94	lavu/mem: un-inline av_size_mult() There seems to be no compelling reason for it to be inline.	3 years ago
Anton Khirnov	c8778606b3	lavu/video_enc_params: make sure blocks are properly aligned	3 years ago
Lynne	08d933bf61	hwcontext_vulkan: fix typo in vulkan_device_init() load_functions() did not load the device-level functions.	3 years ago
Andreas Rheinhardt	7e03d962a4	avutil/opt: Check directly for av_dict_copy() failure av_dict_copy() returned void when this code was written. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@outlook.com>	3 years ago
Jin Bo	fd5fd48659	libavcodec/mips: Fix build errors reported by clang Clang is more strict on the type of asm operands, float or double type variable should use constraint 'f', integer variable should use constraint 'r'. Signed-off-by: Jin Bo <jinbo@loongson.cn> Reviewed-by: yinshiyou-hf@loongson.cn Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>	4 years ago
Valerii Zapodovnikov	6b1268f8c3	pixfmt: fixed wrong fix of comment This mostly reverts `785bfb1d7b`. But I also added some clarifications so that nobody mixes primaries with matrix again. SMPTE 240 and 170 primaires are the same, while matrix coeff. are different, because 240 is derived from 170's new primaries and white point while 170 uses BT.601 derived from BT.470 System M (yes, with Illuminant C) a.k.a. NTSC 1953. Some nits too. Reviewed-by: Reto Kromer <lists@reto.ch> Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>	4 years ago
James Almer	baf5cc5b7a	avutil/mem: use GCC builtins to check for overflow in av_size_mult() Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
James Almer	918fc9a0ed	avutil/mem: check for max_alloc_size in av_fast_malloc() This puts av_fast_malloc*() in line with av_fast_realloc(). Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
James Almer	786be70e28	avutil/mem: make ff_fast_malloc() internal to mem.c Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
James Almer	be96f4b616	avutil/mem: make max_alloc_size an atomic type Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
James Almer	fc99d59553	avutil/imgutils: don't add offsets to NULL pointers Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
Shiyou Yin	ab04fedaaa	mips: Fix potential illegal instruction error. MSA2 optimizations are attached to MSA macros in generic_macros_msa.h. It's difficult to do runtime check for them. Remove this part of code can make it more robust. H264 1080p decoding: 5.13x==>5.12x. Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>	4 years ago
Andreas Rheinhardt	8b83a4a885	avutil/mem: Also poison new av_realloc-allocated blocks Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@outlook.com>	4 years ago
Lynne	cf17e2323f	hwcontext_vulkan: dlopen libvulkan While Vulkan itself went more or less the way it was expected to go, libvulkan didn't quite solve all of the opengl loader issues. It's multi-vendor, yes, but unfortunately, the code is Google/Khronos QUALITY, so suffers from big static linking issues (static linking on anything but OSX is unsupported), has bugs, and due to the prefix system used, there are 3 or so ways to type out functions. Just solve all of those problems by dlopening it. We even have nice emulation for it on Windows.	4 years ago
Lynne	4a6581e968	hwcontext_vulkan: dynamically load functions This patch allows for alternative loader implementations.	4 years ago
James Almer	ffeeff4fbc	avutil/hwcontext_vulkan: fix format specifiers for some printed variables VkPhysicalDeviceLimits.optimalBufferCopyRowPitchAlignment and VkPhysicalDeviceExternalMemoryHostPropertiesEXT.minImportedHostPointerAlignment are of type VkDeviceSize (a typedef uint64_t). VkPhysicalDeviceLimits.minMemoryMapAlignment is of type size_t. Signed-off-by: James Almer <jamrial@gmail.com> Reviewed-by: Lynne <dev@lynne.ee>	4 years ago
Lynne	3a3e8c35b6	hwcontext_vulkan: reorder structure fields and add spaces in between We're in the middle of an ABI unstable period, so we're allowed to.	4 years ago
Anton Khirnov	85ba17f36d	Bump major versions of all libraries. Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@outlook.com>	4 years ago
James Almer	0bf3a7361d	avutil: remove deprecated AVClass.child_class_next Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
Andreas Rheinhardt	d40bb518b5	avutil/cpu: Remove deprecated functions Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@outlook.com> Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
Andreas Rheinhardt	ef6a9e5e31	avutil/buffer: Switch AVBuffer API to size_t Announced in `14040a1d91`. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@outlook.com> Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
Andreas Rheinhardt	985c0dac67	avutil/pixdesc: Remove deprecated AV_PIX_FMT_FLAG_PSEUDOPAL Deprecated in `d6fc031caf`. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
Andreas Rheinhardt	1eb3110115	avutil/frame: Remove deprecated getters and setters Deprecated in `7df37dd319`. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
Andreas Rheinhardt	a240097ecd	avutil: Switch crypto APIs to size_t Announced in `e435beb1ea`. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
Andreas Rheinhardt	6e30b35b85	avutil/frame: Remove deprecated AVFrame.pkt_pts field Deprecated in `32c8359093`. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
Andreas Rheinhardt	3b56fa85e8	avutil/frame: Remove deprecated AVFrame.error Deprecated in `1aa24df74c`. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
Andreas Rheinhardt	0181162bb5	avutil/pixdesc: Remove deprecated off-by-one fields from pix fmt descs Deprecated in `2268db2cd0`. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
Andreas Rheinhardt	b8accd1175	avutil/frame: Remove AVFrame QP table API Originally deprecated in 1296b1f6c0631ab79464e22d48a6a1548450b943; scheduled again for removal in `a991526832`. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
Andreas Rheinhardt	ad524cb9ee	avutil/pixfmt: Remove deprecated VAAPI pixel formats Deprecated in `9f8e57efe4`. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
James Almer	7a6ea6ce2a	x86/tx_float: remove ff_ prefix from external constant tables Fixes compilation with some assemblers. Reviewed-by: Lynne Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
Lynne	bb40f800bd	x86/tx_float: fix forgotten 2-argument mulps Yasm really cannot deal with any omitted arguments at all.	4 years ago
Lynne	e2cf0a1f68	x86/tx_float: use all arguments on vperm2f and vpermilps and reindent comments Apparently even old nasm isn't required to accept incomplete instructions.	4 years ago
James Almer	fddddc7ec2	x86/tx_float: Fixes compilation with old yasm Use three operand format on some instructions, and lea to load effective addresses of tables. Signed-off-by: James Almer <jamrial@gmail.com>	4 years ago
Lynne	e448a4b4ea	lavu/x86/tx_float: fix FMA3 implying AVX2 is available It's the other way around - AVX2 implies FMA3 is available.	4 years ago
Lynne	119a3f7e8d	lavu/x86: add FFT assembly This commit adds a pure x86 assembly SIMD version of the FFT in libavutil/tx. The design of this pure assembly FFT is pretty unconventional. On the lowest level, instead of splitting the complex numbers into real and imaginary parts, we keep complex numbers together but split them in terms of parity. This saves a number of shuffles in each transform, but more importantly, it splits each transform into two independent paths, which we process using separate registers in parallel. This allows us to keep all units saturated and lets us use all available registers to avoid dependencies. Moreover, it allows us to double the granularity of our per-load permutation, skipping many expensive lookups and allowing us to use just 4 loads per register, rather than 8, or in case FMA3 (and by extension, AVX2), use the vgatherdpd instruction, which is at least as fast as 4 separate loads on old hardware, and quite a bit faster on modern CPUs). Higher up, we go for a bottom-up construction of large transforms, foregoing the traditional per-transform call-return recursion chains. Instead, we always start at the bottom-most basis transform (in this case, a 32-point transform), and continue constructing larger and larger transforms until we return to the top-most transform. This way, we only touch the stack 3 times per a complete target transform: once for the 1/2 length transform and two times for the 1/4 length transform. The combination algorithm we use is a standard Split-Radix algorithm, as used in our C code. Although a version with less operations exists (Steven G. Johnson and Matteo Frigo's "A modified split-radix FFT with fewer arithmetic operations", IEEE Trans. Signal Process. 55 (1), 111–119 (2007), which is the one FFTW uses), it only has 2% less operations and requires at least 4x the binary code (due to it needing 4 different paths to do a single transform). That version also has other issues which prevent it from being implemented with SIMD code as efficiently, which makes it lose the marginal gains it offered, and cannot be performed bottom-up, requiring many recursive call-return chains, whose overhead adds up. We go through a lot of effort to minimize load/stores by keeping as much in registers in between construcring transforms. This saves us around 32 cycles, on paper, but in reality a lot more due to load/store aliasing (a load from a memory location cannot be issued while there's a store pending, and there are only so many (2 for Zen 3) load/store units in a CPU). Also, we interleave coefficients during the last stage to save on a store+load per register. Each of the smallest, basis transforms (4, 8 and 16-point in our case) has been extremely optimized. Our 8-point transform is barely 20 instructions in total, beating our old implementation 8-point transform by 1 instruction. Our 2x8-point transform is 23 instructions, beating our old implementation by 6 instruction and needing 50% less cycles. Our 16-point transform's combination code takes slightly more instructions than our old implementation, but makes up for it by requiring a lot less arithmetic operations. Overall, the transform was optimized for the timings of Zen 3, which at the time of writing has the most IPC from all documented CPUs. Shuffles were preferred over arithmetic operations due to their 1/0.5 latency/throughput. On average, this code is 30% faster than our old libavcodec implementation. It's able to trade blows with the previously-untouchable FFTW on small transforms, and due to its tiny size and better prediction, outdoes FFTW on larger transforms by 11% on the largest currently supported size.	4 years ago
Lynne	1978b143eb	checkasm: add av_tx FFT SIMD testing code This sadly required making changes to the code itself, due to the same context needing to be reused for both versions. The lookup table had to be duplicated for both versions.	4 years ago
Lynne	ff71671d88	lavu/tx: add parity revtab generator version This will be used for SIMD support.	4 years ago
Lynne	18af1ea8d1	lavu: bump minor and add APIchanges entry for the lavu/tx changes	4 years ago
Lynne	0072a42388	lavu/tx: add full-sized iMDCT transform flag	4 years ago
Lynne	aa6c757d50	lavu/tx: add unaligned flag to the API	4 years ago
Lynne	8c55c82583	lavu/tx: add a 9-point FFT and (i)MDCT	4 years ago
Lynne	bd9ea917a3	lavu/tx: add a 7-point FFT and (i)MDCT	4 years ago
Lynne	89da62f2fc	lavu/tx: refactor power-of-two FFT This commit refactors the power-of-two FFT, making it faster and halving the size of all tables, making the code much smaller on all systems. This removes the big/small pass split, because on modern systems the "big" pass is always faster, and even on older machines there is no measurable speed difference.	4 years ago
Lynne	aa910a7ecd	lavu/tx: minor code style improvements and additional comments	4 years ago
Andreas Rheinhardt	7368e5537d	avutil/cpu: Schedule deprecated functions for removal av_set_cpu_flags_mask() has been deprecated in the commit which merged it: 6df42f98746be06c883ce683563e07c9a2af983f; av_parse_cpu_flags() has been deprecated in `4b529edff8`. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@outlook.com>	4 years ago
Andreas Rheinhardt	f3c197b129	Include attributes.h directly Some files currently rely on libavutil/cpu.h to include it for them; yet said file won't use include it any more after the currently deprecated functions are removed, so include attributes.h directly. Signed-off-by: Andreas Rheinhardt <andreas.rheinhardt@outlook.com>	4 years ago
Brad Smith	c8fb68ec52	avutil/cpu: Use HW_NCPUONLINE to detect # of online CPUs with OpenBSD Signed-off-by: Brad Smith <brad@comstyle.com> Signed-off-by: Marton Balint <cus@passwd.hu>	4 years ago
Guo, Yejun	0c7aef84a0	lavu/detection_bbox.h: use AV_NUM_DETECTION_BBOX_CLASSIFY to replace AV_NUM_BBOX_CLASSIFY	4 years ago
Lynne	6c65e49990	lavu/detection_bboxes: add missing space Could at least maintainers with push access follow the code styles we have?	4 years ago
Guo, Yejun	f1bf465aa0	lavu: add side data AV_FRAME_DATA_DETECTION_BBOXES for object detection/classification	4 years ago

... 4 5 6 7 8 ...

5504 Commits (bdf01a9609e49ff602b38826420252356d98ba2a)