quakeforge

Commit Graph

Author	SHA1	Message	Date
D G Turner	b799d48ccb	[simd] fix build when avx2 is not available, but avx is. This failed with errors such as: from ./include/QF/simd/vec4d.h:32, from libs/util/simd.c:37: ./include/QF/simd/vec4d.h: In function ‘qmuld’: /usr/lib/gcc/x86_64-pc-linux-gnu/10.3.0/include/avx2intrin.h:1049:1: error: inlining failed in call to ‘always_inline’ ‘_mm256_permute4x64_pd’: target specific option mismatch 1049 \| _mm256_permute4x64_pd (__m256d __X, const int __M)	2021-06-23 01:10:42 +01:00
Bill Currie	93167279fc	Fix a bunch of issues found by gcc-11	2021-06-13 14:30:59 +09:00
Bill Currie	778c07e91f	[util] Get vectors working for non-SSE archs GCC does a fairly nice job of producing code for vector types when the hardware doesn't support SIMD, but it seems to break certain math optimization rules due to excess precision (?). Still, it works well enough for the core engine, but may not be well suited to the tools. However, so far, only qfvis uses vector types (and it's not tested yet), and tools should probably be used on suitable machines anyway (not forces, of course).	2021-06-01 18:53:53 +09:00
Bill Currie	cb7d2671d8	[simd] Make mmulf safe for src=dst I guess I'd forgotten that the parameters are actually pointers.	2021-04-29 19:25:31 +09:00
Bill Currie	e6bc5e3e11	[simd] Add qexpf function	2021-04-25 15:02:08 +09:00
Bill Currie	9ac4cdc6bd	[simd] Fix more portability issues I had missed vec4d.h because it's mostly unused at this stage.	2021-04-02 23:25:14 +09:00
Bill Currie	b6ab832ed4	[simd] Add vabsf and some more tests	2021-03-28 19:49:43 +09:00
Bill Currie	29e029c792	[util] Add float a simd version of the SEB And its support functions. I can't tell if it's any faster (mtwist_rand is a significant chunk of the benchmark timings, oops), but it's nice to have.	2021-03-27 23:38:10 +09:00
Bill Currie	8309e1852a	[simd] Fix some portability issues Use [u]int64_t instead of long, and fix some incorrect attribute usage (I had misread the gcc docs at the time).	2021-03-27 20:04:10 +09:00
Bill Currie	5158cc5527	[util] Add normal and magnitude float vector functions	2021-03-19 11:09:57 +09:00
Bill Currie	5949753579	Make m3vmulf return v[3] unchanged	2021-03-10 19:40:19 +09:00
Bill Currie	09e1a63470	[util] Add a simd mat4 transpose function	2021-03-09 23:50:32 +09:00
Bill Currie	4a97bc3ba5	[util] Create simd quaternion to matrix function This seems to be pretty close to as fast as it gets (might be able to do better with some shuffles of the negation constants instead of loading separate constants).	2021-03-04 17:45:10 +09:00
Bill Currie	45c0255643	[util] Add simd 4x4 matrix functions Currently just add, subtract, multiply (m m and m v).	2021-03-03 16:34:16 +09:00
Bill Currie	9039c6975a	[util] Clean up some missed vsqrt changes	2021-01-05 08:35:53 +09:00
Bill Currie	015cee7b6f	[util] Add vector-quaternion shortcut functions Care needs to be taken to ensure the right function is used with the right arguments, but with these, the need to use qconj(d\|f) for a one-off inverse rotation is removed.	2021-01-02 10:44:45 +09:00
Bill Currie	7bf90e5f4a	[util] Sort out implementation issues for simd	2021-01-02 09:55:59 +09:00
Bill Currie	3125009a7c	[util] Add vector and quaternion types to cexpr Although there's no distinction between the two at the C level, I think it's probably best to separate them in a scripting language.	2020-12-30 18:20:11 +09:00
Bill Currie	1ddd57b09e	[util] Add qconj, vtrunc, vceil and vfloor functions I had forgotten these rather critical functions. Both double and float versions are included.	2020-12-30 18:20:11 +09:00
Bill Currie	09a10f80e1	[util] Add basic SIMD implemented vector functions They take advantage of gcc's vector_size attribute and so only cross, dot, qmul, qvmul and qrot (create rotation quaternion from two vectors) are needed at this stage as basic (per-component) math is supported natively by gcc. The provided functions work on horizontal (array-of-structs) data, ie a vec4d_t or vec4f_t represents a single vector, or traditional vector layout. Vertical layout (struct-of-arrays) does not need any special functions as the regular math can be used to operate on four vectors at a time. Functions are provided for loading a vec4 from a vec3 (4th element set to 0) and storing a vec4 into a vec3 (discarding the 4th element). With this, QF will require AVX2 support (needed for vec4d_t). Without support for doubles, SSE is possible, but may not be worthwhile for horizontal data. Fused-multiply-add is NOT used because it alters the results between unoptimized and optimized code, resulting in -mfma really meaning -mfast-math-anyway. I really do not want to have to debug issues that occur only in optimized code.	2020-12-30 18:20:11 +09:00

20 Commits