llvm-mirror

mirror of https://github.com/RPCS3/llvm-mirror.git synced 2025-01-31 20:51:52 +01:00

Author	SHA1	Message	Date
Nemanja Ivanovic	5779a8862c	[Power9] Exploit move and splat instructions for build_vector improvement This patch corresponds to review: https://reviews.llvm.org/D21135 This patch exploits the following instructions: mtvsrws lxvwsx mtvsrdd mfvsrld In order to improve some build_vector and extractelement patterns. llvm-svn: 282246	2016-09-23 13:25:31 +00:00
James Molloy	2b4c2966f2	[ARM] Promote small global constants to constant pools If a constant is unamed_addr and is only used within one function, we can save on the code size and runtime cost of an indirection by changing the global's storage to inside the constant pool. For example, instead of: ldr r0, .CPI0 bl printf bx lr .CPI0: &format_string format_string: .asciz "hello, world!\n" We can emit: adr r0, .CPI0 bl printf bx lr .CPI0: .asciz "hello, world!\n" This can cause significant code size savings when many small strings are used in one function (4 bytes per string). This recommit contains fixes for a nasty bug related to fast-isel fallback - because fast-isel doesn't know about this optimization, if it runs and emits references to a string that we inline (because fast-isel fell back to SDAG) we will end up with an inlined string and also an out-of-line string, and we won't emit the out-of-line string, causing backend failures. It also contains fixes for emitting .text relocations which made the sanitizer bots unhappy. llvm-svn: 282241	2016-09-23 12:15:58 +00:00
Tom Stellard	a377ff1511	AMDGPU/SI: Include implicit arguments in kernarg_segment_byte_size Reviewers: arsenm Subscribers: arsenm, kzhuravl, wdng, nhaehnle, yaxunl, llvm-commits, tony-tye Differential Revision: https://reviews.llvm.org/D24835 llvm-svn: 282223	2016-09-23 01:33:26 +00:00
Arnold Schwaighofer	115971c664	i386 does not support optimized swifterror handling rdar://28432565 llvm-svn: 282186	2016-09-22 20:06:25 +00:00
Hans Wennborg	d1cad41742	Win64: Don't emit unwind info for "leaf" functions (PR30337) According to MSDN (see the PR), functions which don't touch any callee-saved registers (including %rsp) don't need any unwind info. This patch makes LLVM not emit unwind info for such functions, to save binary size. Differential Revision: https://reviews.llvm.org/D24748 llvm-svn: 282185	2016-09-22 19:50:05 +00:00
Nemanja Ivanovic	07dc6b6937	[PowerPC] Sign extend sub-word values for atomic comparisons Atomic comparison instructions use the sub-word load instruction on Power8 and up but the value is not sign extended prior to the signed word compare instruction. This patch adds that sign extension. llvm-svn: 282182	2016-09-22 19:06:38 +00:00
Nirav Dave	f3f5f12c53	[DAG] Fix incorrect alignment of ext load. Correctly use alignment size from loaded size not output value size. Reviewers: jyknight, tstellarAMD, arsenm Subscribers: llvm-commits Differential Revision: https://reviews.llvm.org/D23356 llvm-svn: 282177	2016-09-22 17:28:43 +00:00
Krzysztof Parzyszek	344cd70c0b	[PPC] Set SP after loading data from stack frame, if no red zone is present Follow-up to r280705: Make sure that the SP is only restored after all data is loaded from the stack frame, if there is no red zone. This completes the fix for https://llvm.org/bugs/show_bug.cgi?id=26519. Differential Revision: https://reviews.llvm.org/D24466 llvm-svn: 282174	2016-09-22 17:22:43 +00:00
Tim Northover	3841e1ae81	GlobalISel: handle stack-based parameters on AArch64. llvm-svn: 282153	2016-09-22 13:49:25 +00:00
Nemanja Ivanovic	a22f68829a	[Power9] Add exploitation of non-permuting memory ops This patch corresponds to review: https://reviews.llvm.org/D19825 The new lxvx/stxvx instructions do not require the swaps to line the elements up correctly. In order to select them over the lxvd2x/lxvw4x instructions which require swaps, the patterns for the old instruction have a predicate that ensures they won't be selected on Power9 and newer CPUs. llvm-svn: 282143	2016-09-22 09:52:19 +00:00
Craig Topper	daf6088084	[AVX-512] Add support for commuting VPTERNLOG instructions. VPTERNLOG is a ternary instruction with an immediate specifying the logical operation to perform. For each bit position in the 3 source vectors the bit from each source is concatenated together and the resulting 3-bit value is used to select a bit in the immediate. This bit value is written to the result vector. We can commute this by swapping operands and modifying the immediate. To modify the immediate we need to swap two pairs of bits. The pairs correspond to the locations in the immediate where the commuted operands bits have opposite values and the uncommuted operand has the same value. Bits 0 and 7 will never be swapped since the relevant bits from all sources are the same value. This refactors and reuses parts of the FMA3 commuting code which is also a three operand instruction. llvm-svn: 282132	2016-09-22 03:00:50 +00:00
Arnold Schwaighofer	f027af8c85	Disable tail calls if there is an swifterror argument ISel does not handle them correctly yet i.e we crash trying to emit tail call code. radar://28407842 llvm-svn: 282088	2016-09-21 16:53:36 +00:00
Cameron McInally	f21a9082e2	[AVX512] Fix return types on int_x86_avx512_gatherXXX_di intrinsics The return type should match the pass through vector type. Differential Revision: https://reviews.llvm.org/D24744 llvm-svn: 282081	2016-09-21 16:06:10 +00:00
Nico Weber	e87f9a86c2	Revert r281715, it caused PR30475 llvm-svn: 282076	2016-09-21 15:33:24 +00:00
Tim Northover	588e105a97	GlobalISel: produce correct code for signext/zeroext ABI flags. We still don't really have an equivalent of "AssertXExt" in DAG, so we don't exploit the guarantees on the receiving side yet, but this should produce conservatively correct code on iOS ABIs. llvm-svn: 282069	2016-09-21 12:57:45 +00:00
Simon Dardis	55f9daf6b9	[mips] LLVM PR/30197 - Tail call incorrectly clobbers arguments for mips The postRA scheduler performs alias analysis to determine if stores and loads can moved past each other. When a function has more arguments than argument registers for the calling convention used, excess arguments are spilled onto the stack. LLVM by default assumes that argument slots are immutable, unless the function contains a tail call. Without the knowledge of that a function contains a tail call site, stores and loads to fixed stack slots may be re-ordered causing the out-going arguments to clobber the incoming arguments before the incoming arguments are supposed to be dead. Reviewers: vkalintiris Differential Review: https://reviews.llvm.org/D24077 llvm-svn: 282063	2016-09-21 09:43:40 +00:00
NAKAMURA Takumi	675f769e36	llvm/test/CodeGen/NVPTX/zero-cs.ll: Relax an expression to match in -Asserts. LLVM ERROR: Cannot select: 0x3607bf0: i32 = ExternalSymbol'__powidf2' llvm-svn: 282053	2016-09-21 04:43:11 +00:00
Jacques Pienaar	92806722a6	[NVPTX] Check if callsite is defined when computing argument allignment Summary: In getArgumentAlignment check if the ImmutableCallSite pointer CS is non-null before dereferencing. If CS is 0x0 fall back to the ABI type alignment else compute the alignment as before. Reviewers: eliben, jpienaar Subscribers: jlebar, vchuravy, cfe-commits, jholewinski Differential Revision: https://reviews.llvm.org/D9168 llvm-svn: 282045	2016-09-21 01:57:57 +00:00
Petr Hosek	658d5cf8a0	Mark ELF sections whose name start with .note as note Previously, such section would be marked as SHT_PROGBITS which makes it impossible to use an initialized C variable declaration to emit an (allocated) ELF note. The new behavior is also consistent with ELF assembly parser. Differential Revision: https://reviews.llvm.org/D24692 llvm-svn: 282010	2016-09-20 20:21:13 +00:00
Sanjay Patel	d54b3459fa	[x86] split up tests, regenerate checks Note that we fail to eliminate 'or' with 0! llvm-svn: 282005	2016-09-20 19:31:30 +00:00
Evandro Menezes	1c47b35a01	Revert "[AArch64] Use the reciprocal estimation machinery" This reverts commit b7d42b0048f65346e9fa37fb65defeea7ce8c337 per request by Eric Christopher <echristo@gmail.com> (v. http://bit.ly/2cmz6kW). llvm-svn: 282000	2016-09-20 19:02:06 +00:00
Evandro Menezes	27cb59936a	Revert "[AArch64] Properly validate the reciprocal estimation." This reverts commit ad8ca1528242e2a4cb363e3779309e70eb7a430e per request by Eric Christopher <echristo@gmail.com> (v. http://bit.ly/2cmz6kW). llvm-svn: 281999	2016-09-20 19:02:02 +00:00
Tim Northover	c5866724c9	GlobalISel: split aggregates for PCS lowering This should match the existing behaviour for passing complicated struct and array types, in particular HFAs come through like that from Clang. For C & C++ we still need to somehow support all the weird ABI flags, or at least those that are present in the IR (signext, byval, ...), and stack-based parameter passing. llvm-svn: 281977	2016-09-20 15:20:36 +00:00
Simon Pilgrim	1446648b40	[X86][SSE] Regenerate multiple combine tests llvm-svn: 281973	2016-09-20 14:42:45 +00:00
Elena Demikhovsky	943a1dadb1	AVX-512: Fixed a bug in lowering saturated operations on KNL. The generated code is still not optimal. Differential Revision: https://reviews.llvm.org/D24723 llvm-svn: 281966	2016-09-20 11:02:26 +00:00
Craig Topper	6651d4f318	[AVX-512] Use 512-bit vcvtps2ph/vcvtph2ps to implement fp_to_f16/f16_to_fp when F16C and VLX are not supported. Fixes PR23941. llvm-svn: 281958	2016-09-20 05:44:47 +00:00
Matthias Braun	140d944d08	BranchFolder: Fix invalid undef flags after merge. It is legal to merge instructions with different undef flags; However we must drop the undef flag from the merged instruction if it isn't present everywhere. This fixes http://llvm.org/PR30199 llvm-svn: 281957	2016-09-20 01:14:42 +00:00
Sanjay Patel	f78534612c	[x86] auto-generate checks llvm-svn: 281950	2016-09-19 23:44:50 +00:00
Simon Pilgrim	bb3da42dcc	[X86][SSE] Updated vector abs tests Renamed and added v2i64 / v4i64 tests llvm-svn: 281937	2016-09-19 20:50:35 +00:00
Matthias Braun	c02cdc0d60	LiveRangeCalc: Fix reporting of invalid vreg usage in liveness calculation Machine programs need a definition of each vreg before reaching a use (the definition may come from an IMPLICIT_DEF instruction). This class of errors is not detected by the MachineVerifier because of efficiency concerns. LiveRangeCalc used to report these problems, make it do that again (followup to r279625). Also use report_fatal_error() instead of llvm_unreachable() as the error reporting is only present in asserts build anyway. llvm-svn: 281914	2016-09-19 16:49:45 +00:00
Nico Weber	99a8836600	Revert r281841, it does not work on Windows (PR30443). llvm-svn: 281905	2016-09-19 15:22:04 +00:00
Tim Northover	11b2b57958	ARM: check alignment before transforming ldr -> ldm (or similar). ldm and stm instructions always require 4-byte alignment on the pointer, but we weren't checking this before trying to reduce code-size by replacing a post-indexed load/store with them. Unfortunately, we were also dropping this incormation in DAG ISel too, but that's easy enough to fix. llvm-svn: 281893	2016-09-19 09:11:09 +00:00
Elena Demikhovsky	53520a83c1	[X86 Codegen Test] Divided masked_memop into several files. NFC. The masked_memop.ll became huge. I extracted AVX-512 specific tests into separate files. llvm-svn: 281892	2016-09-19 08:58:43 +00:00
Craig Topper	c7f779bf05	[X86,AVX-512] Use INSERT_SUBREG instead of SUBREG_TO_REG when the input is not the output of an instruction. SUBREG_TO_REG is supposed to indicate that the super register has been zeroed, but we can't prove that if we don't know where it came from. llvm-svn: 281885	2016-09-19 02:53:43 +00:00
Craig Topper	9f5d417d08	[AVX-512] Add support for lowering fp_to_f16 and f16_to_fp when VLX is supported regardless of whether F16C is also supported. Still need to add support for lowering using AVX512F when neither VLX or F16C is supported. llvm-svn: 281884	2016-09-19 02:53:37 +00:00
Dean Michael Berris	fca2ef5fb0	[XRay] ARM 32-bit no-Thumb support in LLVM This is a port of XRay to ARM 32-bit, without Thumb support yet. The XRay instrumentation support is moving up to AsmPrinter. This is one of 3 commits to different repositories of XRay ARM port. The other 2 are: https://reviews.llvm.org/D23932 (Clang test) https://reviews.llvm.org/D23933 (compiler-rt) Differential Revision: https://reviews.llvm.org/D23931 llvm-svn: 281878	2016-09-19 00:54:35 +00:00
Simon Pilgrim	b15b82c418	[X86][SSE] Improve recognition of uitofp conversions that can be performed as sitofp With D24253 we can now use SelectionDAG::SignBitIsZero with vector operations. This patch uses SelectionDAG::SignBitIsZero to recognise that a zero sign bit means that we can use a sitofp instead of a uitofp (which is not directly support on pre-AVX512 hardware). While AVX512 does provide support for uitofp, the conversion to sitofp should not cause any regressions. Differential Revision: https://reviews.llvm.org/D24343 llvm-svn: 281852	2016-09-18 12:45:23 +00:00
Wei Mi	4f31f47bed	Change the order of the splitted store from high - low to low - high. It is a trivial change which could make the testcase easier to be reused for the store splitting in CodeGenPrepare. llvm-svn: 281846	2016-09-18 06:10:32 +00:00
Simon Pilgrim	02f84f8e9d	[X86][SSE] Added vector udiv combine tests llvm-svn: 281842	2016-09-17 22:02:23 +00:00
Simon Pilgrim	08eae71db6	[X86][SSE] Added vector fcopysign combine tests Also demonstrating the poor lowering of fcopysign... llvm-svn: 281841	2016-09-17 21:31:34 +00:00
Simon Pilgrim	535a47bf28	[X86][SSE] Added vector mul combine tests llvm-svn: 281839	2016-09-17 20:06:16 +00:00
Simon Pilgrim	8b1084e1d0	[X86][SSE] Improve target shuffle mask extraction Add ability to extract vXi64 'vzext_movl' masks on 32-bit targets llvm-svn: 281834	2016-09-17 18:50:54 +00:00
Simon Pilgrim	466154ecbb	[X86][AVX] Test target shuffle combining on 32 and 64-bit targets llvm-svn: 281833	2016-09-17 18:42:41 +00:00
Simon Pilgrim	75052d70a3	[X86][AVX2] Add target shuffle constant folding tests llvm-svn: 281830	2016-09-17 17:42:15 +00:00
Simon Pilgrim	b7d2d32e7a	[X86][AVX] Add target shuffle constant folding tests llvm-svn: 281829	2016-09-17 17:41:14 +00:00
Simon Pilgrim	35f14c5775	[X86][XOP] Add target shuffle constant folding tests llvm-svn: 281828	2016-09-17 17:40:40 +00:00
Simon Pilgrim	b44df8e06b	[X86][SSSE3] Add target shuffle constant folding tests llvm-svn: 281827	2016-09-17 17:40:08 +00:00
Ron Lieberman	374b9a54f7	[Hexagon] segv while processing SUnit with nullNodePtr Added BoundaryNode check to isBestZeroLatency function. llvm-svn: 281825	2016-09-17 16:21:09 +00:00
Matt Arsenault	17a8bb755d	AMDGPU: Fix broken FrameIndex handling We were trying to avoid using a FrameIndex operand in non-pointer operands in a convoluted way, and would break because of using TargetFrameIndex. The TargetFrameIndex should only be used in the case where it makes sense to fold it as part of the addressing mode, otherwise it requires materialization like a normal constant. This wasn't working reliably and failed in the added testcase, hitting the assert when processing the frame index. The TargetFrameIndex was coming from trying to produce an AssertZext limiting the maximum stack size. I'm not sure this was correct to begin with, because it is apparently possible to have a single workitem dispatch that requires all 4G of private memory. llvm-svn: 281824	2016-09-17 16:09:55 +00:00
Matt Arsenault	8c9bd8e031	AMDGPU: Push bitcasts through build_vector This reduces the number of copies and reg_sequences when using fp constant vectors. This significantly reduces the code size in local-stack-alloc-bug.ll llvm-svn: 281822	2016-09-17 15:44:16 +00:00

1 2 3 4 5 ...

17522 Commits