Ronald S. Bultje
6341838f3c
Use word-writing instead of dword-writing (with two cached but otherwise
...
unchanged bytes) in the horizontal simple loopfilter. This makes the filter
quite a bit faster in itself (~30 cycles less on Core1), probably mostly
because we don't need a complex 4x4 transpose, but only a simple byte
interleave. Also allows using pextrw on SSE4, which speeds up even more
(e.g. 25% faster on Core i7).
Originally committed as revision 24638 to svn://svn.ffmpeg.org/ffmpeg/trunk
2010-07-31 23:13:15 +00:00
..
2010-07-31 23:13:15 +00:00
2010-07-24 17:11:51 +00:00
2010-07-24 02:57:08 +00:00
2010-07-24 13:59:49 +00:00
2010-07-25 14:33:16 +00:00
2010-07-24 13:59:49 +00:00
2010-07-27 07:18:36 +00:00
2010-07-27 15:54:26 +00:00
2010-07-29 23:44:57 +00:00
2010-07-31 22:15:59 +00:00
2010-07-27 21:12:16 +00:00
2010-07-31 21:14:03 +00:00
2010-07-24 13:59:49 +00:00
2010-07-27 17:11:13 +00:00
2010-07-23 00:34:09 +00:00
2010-07-31 16:46:20 +00:00
2010-07-31 16:46:20 +00:00
2010-07-28 08:02:35 +00:00
2010-07-27 07:18:36 +00:00
2010-07-27 15:54:26 +00:00
2010-07-27 07:18:36 +00:00
2010-07-27 15:54:26 +00:00
2010-07-27 10:08:34 +00:00
2010-07-24 13:59:49 +00:00
2010-07-28 05:19:42 +00:00
2010-07-24 13:59:49 +00:00
2010-07-24 13:59:49 +00:00
2010-07-31 16:46:20 +00:00
2010-07-24 13:59:49 +00:00
2010-07-28 05:36:33 +00:00
2010-07-28 05:38:30 +00:00
2010-07-27 23:09:13 +00:00
2010-07-23 21:46:25 +00:00
2010-07-23 06:02:52 +00:00
2010-07-23 06:02:52 +00:00
2010-07-28 05:40:38 +00:00
2010-07-28 05:38:30 +00:00
2010-07-24 13:59:49 +00:00
2010-07-27 15:54:26 +00:00
2010-07-26 13:52:49 +00:00