FFmpeg

Commit Graph

Author	SHA1	Message	Date
Rostislav Pehlivanov	70eb77b34e	mdct15: add inverse transform postrotation SIMD 2.5ms frames: Before (c): 2638 decicycles in postrotate, 2097040 runs, 112 skips After (sse3): 1467 decicycles in postrotate, 2097083 runs, 69 skips After (avx2): 1244 decicycles in postrotate, 2097085 runs, 67 skips 5ms frames: Before (c): 4987 decicycles in postrotate, 1048371 runs, 205 skips After (sse3): 2644 decicycles in postrotate, 1048509 runs, 67 skips After (avx2): 2031 decicycles in postrotate, 1048523 runs, 53 skips 10ms frames: Before (c): 9153 decicycles in postrotate, 523575 runs, 713 skips After (sse3): 5110 decicycles in postrotate, 523726 runs, 562 skips After (avx2): 3738 decicycles in postrotate, 524223 runs, 65 skips 20ms frames: Before (c): 17857 decicycles in postrotate, 261866 runs, 278 skips After (sse3): 10041 decicycles in postrotate, 261746 runs, 398 skips After (avx2): 7050 decicycles in postrotate, 262116 runs, 28 skips Improves total decoding performance for real world content by 9% with avx2. Signed-off-by: Rostislav Pehlivanov <atomnuker@gmail.com>	7 years ago
Rostislav Pehlivanov	aef5f9ab05	mdct15: remove redundant scale argument to imdct_half The only use of that argument was for Opus downmixing which is very rare and better done after the mdcts. Signed-off-by: Rostislav Pehlivanov <atomnuker@gmail.com>	7 years ago
Rostislav Pehlivanov	e1120b1c54	mdct15: add assembly optimizations for the 15-point FFT c: 1802 decicycles in fft15,16774635 runs, 2581 skips avx: 865 decicycles in fft15,16776378 runs, 838 skips Signed-off-by: Rostislav Pehlivanov <atomnuker@gmail.com>	8 years ago
Rostislav Pehlivanov	d2119f624d	imdct15: rename to mdct15 and add a forward transform Handles strides (needed for Opus transients), does pre-reindexing and folding without needing a copy. Signed-off-by: Rostislav Pehlivanov <atomnuker@gmail.com>	8 years ago
Rostislav Pehlivanov	2d208aaabe	imdct15: replace the FFT with a faster PFA FFT algorithm This commit replaces the current inefficient non-power-of-two FFT with a much faster FFT based on the Prime Factor Algorithm. Although it is already much faster than the old algorithm without SIMD, the new algorithm makes use of the already very throughouly SIMD'd power of two FFT, which improves performance even more across all platforms which we have SIMD support for. Most of the work was done by Peter Barfuss, who passed the code to me to implement into the iMDCT and the current codebase. The code for a 5-point and 15-point FFT was derived from the previous implementation, although it was optimized and simplified, which will make its future SIMD easier. The 15-point FFT is currently using 6% of the current overall decoder overhead. The FFT can now easily be used as a forward transform by simply not multiplying the 5-point FFT's imaginary component by -1 (which comes from the fact that changing the complex exponential's angle by -1 also changes the output by that) and by multiplying the "theta" angle of the main exptab by -1. Hence the deliberately left multiplication by -1 at the end. FATE passes, and performance reports on other platforms/CPUs are welcome. Performance comparisons: iMDCT, PFA: 101127 decicycles in speed, 32765 runs, 3 skips iMDCT, Old: 211022 decicycles in speed, 32768 runs, 0 skips Standalone FFT, 300000 transforms of size 960: PFA Old FFT kiss_fft libfftw3f 3.659695s, 15.726912s, 13.300789s, 1.182222s Being only 3x slower than libfftw3f is a big achievement by itself. There appears to be something capping the performance in the iMDCT side of things, possibly during the pre-stage reindexing. However, it is certainly fast enough for now. Signed-off-by: Rostislav Pehlivanov <atomnuker@gmail.com>	8 years ago
Rostislav Pehlivanov	4fdacf4cdb	imdct15: remove the AArch64 assembly Prep work for the next commit, which will add a new FFT algorithm which makes the iMDCT over 3x faster than it is currently (standalone, the FFT is with some framesizes over 10x faster). The new FFT algorithm uses the already thouroughly SIMD'd power of two FFT which already has SIMD for AArch64, so users of that platform will still see an improvement. The previous FFT+SIMD was barely 2.5x faster than the C versions on these platforms. Signed-off-by: Rostislav Pehlivanov <atomnuker@gmail.com>	8 years ago
Diego Biurrun	3d5d46233c	opus: Factor out imdct15 into a standalone component It will be reused by the AAC decoder.	10 years ago
Janne Grunau	d3f5b94762	aarch64: opus NEON iMDCT and FFT Opus celt decoding 11% faster and the iMDCT over 2.5 times faster on Apple's A7.	11 years ago
Diego Biurrun	322a1dda97	dsputil: Refactor duplicated CALL_2X_PIXELS / PIXELS16 macros	11 years ago
Diego Biurrun	acd2b8e42d	rnd_avg.h: K&R formatting cosmetics	11 years ago
Ronald S. Bultje	68d8238cca	hpeldsp: Add half-pel functions (currently copies of dsputil) Signed-off-by: Martin Storsjö <martin@martin.st>	12 years ago
Ronald S. Bultje	9628e5a4ac	hpeldsp: add half-pel functions (currently copies of dsputil).	12 years ago
Michael Niedermayer	5fd6d85d17	rnd_avg: fix author attribution Reference: commit `41fda91d09` Author: BERO <bero@geocities.co.jp> Date: Wed May 14 17:46:55 2003 +0000 aligned dsputil (for sh4) patch by (BERO <bero at geocities dot co dot jp>) Originally committed as revision 1880 to svn://svn.ffmpeg.org/ffmpeg/trunk commit `8dbe585641` Author: Oskar Arvidsson <oskar@irock.se> Date: Tue Mar 29 17:48:59 2011 +0200 Adds 8-, 9- and 10-bit versions of some of the functions used by the h264 decoder. This patch lets e.g. dsputil_init chose dsp functions with respect to the bit depth to decode. The naming scheme of bit depth dependent functions is <base name>_<bit depth>[_<prefix>] (i.e. the old clear_blocks_c is now named clear_blocks_8_c). Note: Some of the functions for high bit depth is not dependent on the bit depth, but only on the pixel size. This leaves some room for optimizing binary size. Preparatory patch for high bit depth h264 decoding support. Signed-off-by: Michael Niedermayer <michaelni@gmx.at> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	12 years ago
Diego Biurrun	bf6b3ec924	dsputil: Move rnd_avg inline functions to a separate header	12 years ago
Mans Rullgard	2912e87a6c	Replace FFmpeg with Libav in licence headers Signed-off-by: Mans Rullgard <mans@mansr.com>	14 years ago
Måns Rullgård	8fc0162ac4	Add av_ prefix to bswap macros Originally committed as revision 24170 to svn://svn.ffmpeg.org/ffmpeg/trunk	15 years ago
Diego Biurrun	ba87f0801d	Remove explicit filename from Doxygen @file commands. Passing an explicit filename to this command is only necessary if the documentation in the @file block refers to a file different from the one the block resides in. Originally committed as revision 22921 to svn://svn.ffmpeg.org/ffmpeg/trunk	15 years ago
Måns Rullgård	2ed6f39944	Replace many includes of libavutil/common.h with what is actually needed This reduces the number of false dependencies on header files and speeds up compilation. Originally committed as revision 22407 to svn://svn.ffmpeg.org/ffmpeg/trunk	15 years ago
Diego Biurrun	bad5537e2c	Use full internal pathname in doxygen @file directives. Otherwise doxygen complains about ambiguous filenames when files exist under the same name in different subdirectories. Originally committed as revision 16912 to svn://svn.ffmpeg.org/ffmpeg/trunk	16 years ago
Måns Rullgård	3a90480ac4	split bswap.h into per-arch files Originally committed as revision 15663 to svn://svn.ffmpeg.org/ffmpeg/trunk	16 years ago
Stefano Sabatini	987903826b	Globally rename the header inclusion guard names. Consistently apply this rule: the guard name is obtained from the filename by stripping the leading "lib", converting '/' and '.' to '_' and uppercasing the resulting name. Guard names in the root directory have to be prefixed by "FFMPEG_". Originally committed as revision 15120 to svn://svn.ffmpeg.org/ffmpeg/trunk	16 years ago
Diego Biurrun	245976da2a	Use full path for #includes from another directory. Originally committed as revision 13098 to svn://svn.ffmpeg.org/ffmpeg/trunk	17 years ago
Diego Biurrun	5b21bdabe4	Add FFMPEG_ prefix to all multiple inclusion guards. Originally committed as revision 10765 to svn://svn.ffmpeg.org/ffmpeg/trunk	17 years ago
Aurelien Jacobs	ca6e50afc1	add a ff_ prefix to some mpegaudio funcs Originally committed as revision 9081 to svn://svn.ffmpeg.org/ffmpeg/trunk	18 years ago
Aurelien Jacobs	4bd8e17c8d	loosen dependencies over mpegaudiodec Originally committed as revision 9080 to svn://svn.ffmpeg.org/ffmpeg/trunk	18 years ago
Diego Biurrun	b78e7197a8	Change license headers to say 'FFmpeg' instead of 'this program/this library' and fix GPL/LGPL version mismatches. Originally committed as revision 6577 to svn://svn.ffmpeg.org/ffmpeg/trunk	18 years ago
Luca Barbato	99aed7c8fc	New single instruction math operation header Originally committed as revision 6291 to svn://svn.ffmpeg.org/ffmpeg/trunk	18 years ago
Diego Biurrun	5509bffa88	Update licensing information: The FSF changed postal address. Originally committed as revision 4842 to svn://svn.ffmpeg.org/ffmpeg/trunk	19 years ago
Michael Niedermayer	3f3f8b2b75	cleanup Originally committed as revision 4312 to svn://svn.ffmpeg.org/ffmpeg/trunk	20 years ago
Bernhard Rosenkränzer	6ad1fa5a49	Better ARM support for mplayer/ffmpeg, ported from atty fork while playing with some new hardware, I found it's running a forked mplayer -- and it looks like they're following the GPL. The maintainer's page is here: http://atty.jp/?Zaurus/mplayer Unfortunately it's mostly in Japanese, so it's hard to figure out any details. Their code looks quite interesting (at least to those of us w/ ARM CPUs). The patches I've attached are the patches from atty.jp with a couple of modifications by myself: - ported to current CVS - reverted their change of removing SNOW support from ffmpeg - cleaned up their bswap mess - removed DOS-style linebreaks from various files patch by (Bernhard Rosenkraenzer: bero, arklinux org) Originally committed as revision 4311 to svn://svn.ffmpeg.org/ffmpeg/trunk	20 years ago
Michael Niedermayer	b0368839ac	MpegEncContext.(i)dct_* -> DspContext.(i)dct_* bitexact cleanup Originally committed as revision 1617 to svn://svn.ffmpeg.org/ffmpeg/trunk	22 years ago
Zdenek Kabelac	0c1a9edad4	* UINTX -> uintx_t INTX -> intx_t Originally committed as revision 1578 to svn://svn.ffmpeg.org/ffmpeg/trunk	22 years ago
Zdenek Kabelac	bb28568364	* cut&paste fix Originally committed as revision 1249 to svn://svn.ffmpeg.org/ffmpeg/trunk	22 years ago
Zdenek Kabelac	5940262772	* oops fixed bad initialization of ff vals. - put FF_LIBMPEG2_IDCT_PERM into CVS - so it will work for now Originally committed as revision 1227 to svn://svn.ffmpeg.org/ffmpeg/trunk	22 years ago
Zdenek Kabelac	83f238cbf0	* compilation fix (ARM users please check) Originally committed as revision 1225 to svn://svn.ffmpeg.org/ffmpeg/trunk	22 years ago
Michael Niedermayer	50eb9cbc44	idct_permutation_type variable, so the permutation type can quickly be identified Originally committed as revision 1071 to svn://svn.ffmpeg.org/ffmpeg/trunk	22 years ago
Michael Niedermayer	676e200cff	trying to fix the non-x86 IDCTs (untested) Originally committed as revision 1006 to svn://svn.ffmpeg.org/ffmpeg/trunk	22 years ago
Fabrice Bellard	ff4ec49e64	license/copyright change Originally committed as revision 599 to svn://svn.ffmpeg.org/ffmpeg/trunk	23 years ago
Fabrice Bellard	92651f67a0	arm specific code Originally committed as revision 79 to svn://svn.ffmpeg.org/ffmpeg/trunk	24 years ago

4 Commits (d94dda742c8eab3141197270fb78063ed22442aa)