From patchwork Wed Jul 8 14:46:31 2026 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Andrew Stubbs X-Patchwork-Id: 138755 Return-Path: X-Original-To: patchwork@sourceware.org Delivered-To: patchwork@sourceware.org Received: from vm01.sourceware.org (localhost [IPv6:::1]) by sourceware.org (Postfix) with ESMTP id 0051D4BA2E05 for ; Wed, 8 Jul 2026 14:47:33 +0000 (GMT) DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org 0051D4BA2E05 Authentication-Results: sourceware.org; dkim=pass (2048-bit key, secure) header.d=baylibre.com header.i=@baylibre.com header.a=rsa-sha256 header.s=google header.b=T9+JwCvU X-Original-To: gcc-patches@gcc.gnu.org Delivered-To: gcc-patches@gcc.gnu.org Received: from mail-wr1-x435.google.com (mail-wr1-x435.google.com [IPv6:2a00:1450:4864:20::435]) by sourceware.org (Postfix) with ESMTPS id 3DFCB4BA2E04 for ; Wed, 8 Jul 2026 14:46:47 +0000 (GMT) DMARC-Filter: OpenDMARC Filter v1.4.2 sourceware.org 3DFCB4BA2E04 Authentication-Results: sourceware.org; dmarc=none (p=none dis=none) header.from=baylibre.com Authentication-Results: sourceware.org; spf=pass smtp.mailfrom=baylibre.com ARC-Filter: OpenARC Filter v1.0.0 sourceware.org 3DFCB4BA2E04 Authentication-Results: sourceware.org; arc=none smtp.remote-ip=2a00:1450:4864:20::435 ARC-Seal: i=1; a=rsa-sha256; d=sourceware.org; s=key; t=1783522007; cv=none; b=RcjOQZLtaRvth7mUOCxWRPKNkTkULMs2L6pDG0wrqQWvy63aPcCa6rGe454+PIbasvrxui1S/p3o97Sz9Ewzb0LJh/j+xgoeqeP+bDzs5IybDzzTRSjDaachQm+Oiyv6OkNttIFSIJdJTU4+a3h/ZnArywI2d4AP9c2vWaF7KDo= ARC-Message-Signature: i=1; a=rsa-sha256; d=sourceware.org; s=key; t=1783522007; c=relaxed/simple; bh=DaZkpc/guiT/g+KiEdFD5Ik4wrfWgnIH+NB0wtyqiKw=; h=DKIM-Signature:From:To:Subject:Date:Message-ID:MIME-Version; b=A5X54PIABCteaXkFq30oiG9ZUH12knEM5yD59vBbpyNsK3r3t0L/c+yaR0J8Msh0m//MDsIh9NZylEIijJ2PGLkojjsWvdXTXmfmhPLoR4a6Y0nIe4R+KYZUxmj/OSPGJkRiXyJcYrGTB9HQ+7/Q3ZxKir/clRHLmrud6qoDy88= ARC-Authentication-Results: i=1; sourceware.org; dkim=pass (2048-bit key, secure) header.d=baylibre.com header.i=@baylibre.com header.a=rsa-sha256 header.s=google header.b=T9+JwCvU DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org 3DFCB4BA2E04 Received: by mail-wr1-x435.google.com with SMTP id ffacd0b85a97d-476a130c138so864635f8f.0 for ; Wed, 08 Jul 2026 07:46:47 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=baylibre.com; s=google; t=1783522006; x=1784126806; darn=gcc.gnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:to:from:from:to:cc:subject:date:message-id :reply-to:content-type; bh=wdPK12LCS/Hn8EEJGq1IaP4C08EnO5dS/gP6NL/VfYI=; b=T9+JwCvUSnNxJzmSjMtB7W7dSJwwz5EplXGEbpESM1nJj0QZlWTFdznm2gIYnuwlvE G/YVCAt4M1GKoxB7MFbqgmqRtbhGaIHFXGoKd1yjnRrMgJgqGuRnUgXa79GsEElUvDdh hpyiPF/AAYNuvUdUlhdzvJTBgKQ5hcPmZnz3fBNBNR6CGEA/4dnEqByYCQgQpfl6lpOJ edVkQzCvpc6zoI0KDQ7C6wcjdy2IDdgsYcJZbA8FqbALpeFp+84UfSBJJ04ftEF3NPua Ie2XpGhxIxqRReiK0rMCYOgvN6nYtOw2tJlCTzDsE9iAw6SuEVm9bilYB9mY4386sac4 xBHA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783522006; x=1784126806; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:to:from:x-gm-gg:x-gm-message-state:from:to :cc:subject:date:message-id:reply-to:content-type; bh=wdPK12LCS/Hn8EEJGq1IaP4C08EnO5dS/gP6NL/VfYI=; b=mS/+uFYpPYmkrhYaQ2t+4asgTEDLpiaWVmyBwP72XQT7/9WY0OBMmUyC5KoniZop5W 6osyiC8t9OfQAE7mkQpeS6qGBJ9AL1FosR2AAfNfU3WOj/PPCDGdD63LqgIibb8YnHqE nG3dLdBaxDecKqP5wCYOc5tc6lW2TWlzRVUYLySGWwMyTXW8EUrv5CJKo3IJhf4d6Lg8 0LIqJSBNmKsaw3ZHBQz6xWybDMl6Z8eD6zUM1yuwKjvNs94FNssBHf63g9SFi9Fjmfop rX8yW8eH4zkR1JT0JWZLf2zK0fonHF/7n8QAl1qppBMWbnZ+N6hf9R9CoNpJeLBXJgjU 9K6A== X-Gm-Message-State: AOJu0Yyfa1/NjzltsXn3CGdLB+UPQFW7MlwPqevodq6T5K+qlYXJOBed j0+jTetRH+05zcofJwIL3krIM3992SHk12UbDI308cKQPUCHbfheWhhrs9Pxj0VwOtNEnD+lSre RepxY X-Gm-Gg: AfdE7cmnS63c6m95+xByPXLovVBbt2KUvFsxuf2YGBtD57MjOx+58M9v5pILwYQJTyN ymWf2nFPcU5yJ2/RA/gOi/4kGVrWJwHmtn2yOYljYRSZ4ePKldppbMPYWyfwMT+UD1zNPCl4tUp 5dP56/2BKwgpT+0wEDQ7b+I4gHJfhqOU8sB8tDEjkSt2826RWDiUqjE1rVRVOMinuA8LIYFK0yG PZ3/b5uBdY0v3gHCKxqgCj3B0vSZpQXlgO3hzDnppXWECwvXmCfLu7PHdNa46HKnf/0Y8I0xf0g wJ3P1S3Ij0GKpnvRvD/adleKDGeOYmx1TD9lH2YzGHMIIBYfW0BEuFhm1IVToHoErOn6GWLFZMm ydW9sLG+FtJLaOdrKT6Wj48Gf4iQ2+oUbz5etd3uw++69iLVGn4efG89H+lFuHTSyghpVqpvpGw mXKiFOJRxRZWSmYdg= X-Received: by 2002:a05:6000:3c6:b0:476:7036:f854 with SMTP id ffacd0b85a97d-47df072bac3mr3469837f8f.21.1783522005897; Wed, 08 Jul 2026 07:46:45 -0700 (PDT) Received: from vbuild-02.baylibre ([217.13.61.132]) by smtp.googlemail.com with ESMTPSA id ffacd0b85a97d-47a9e4d83bdsm43362109f8f.13.2026.07.08.07.46.45 for (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 08 Jul 2026 07:46:45 -0700 (PDT) From: Andrew Stubbs To: gcc-patches@gcc.gnu.org Subject: [PATCH 1/3] rtl: Allow "(mem: (reg:))" Date: Wed, 8 Jul 2026 14:46:31 +0000 Message-ID: <20260708144633.1530935-2-ams@baylibre.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260708144633.1530935-1-ams@baylibre.com> References: <20260708144633.1530935-1-ams@baylibre.com> MIME-Version: 1.0 X-Spam-Status: No, score=-11.2 required=5.0 tests=BAYES_00, DKIM_SIGNED, DKIM_VALID, DKIM_VALID_AU, DKIM_VALID_EF, GIT_PATCH_0, RCVD_IN_DNSWL_NONE, SPF_HELO_NONE, SPF_PASS, TXREP shortcircuit=no autolearn=ham autolearn_force=no version=3.4.6 X-Spam-Checker-Version: SpamAssassin 3.4.6 (2021-04-09) on sourceware.org X-BeenThere: gcc-patches@gcc.gnu.org X-Mailman-Version: 2.1.30 Precedence: list List-Id: Gcc-patches mailing list List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: gcc-patches-bounces~patchwork=sourceware.org@gcc.gnu.org This patch makes the middle-end adjustments needed to allow a MEM to take a vector of addresses in order to read/write a non-contiguous vector of data. Mostly the changes don't have to do anything besides letting the vectors pass. These changes, alone, are not enough to actually allow such MEMs, since it's unlikely that any back-end reports the vector modes "legitimate", so there should be no change of behaviour on any existing target (unconfirmed). I've modified the type of only the TARGET_ADDR_SPACE_VALID_POINTER_MODE hook, which is currently not being used by any target (amdgcn will use it in the following patch), and left the TARGET_VALID_POINTER_MODE alone so as not to have to touch all the other back-ends. If another back-end wants to use vectors of addresses then I propose that either TARGET_VALID_POINTER_MODE can be fixed up at that time, or that backend simply switch to using TARGET_ADDR_SPACE_VALID_POINTER_MODE. I have not made any changes to LRA code. I believe this is not necessary as long as the target uses "define_special_memory_constraint". I failed to figure out how to make the regular memory constraints work, and as it's not necessary for my use case I'm loath to invest any more time into that. Hopefully the documentation patches should describe the intended semantics. I've tried to keep it very simple. There's one "TODO" in the code; I'm intentionally applying the principle of "don't add it if you don't need it" here (not least because I don't know how to test it). There's an assert to make sure it never bites anyone silently. gcc/ChangeLog: * doc/rtl.texi: Document new MEM semantics. * doc/tm.texi: Likewise. * emit-rtl.cc (gen_int_mode): Extend to DTRT for vector modes. (adjust_address_1): Permit vectors of addresses. * explow.cc (convert_memory_address_addr_space_1): Permit vectors of addresses, for no-op conversions only. * recog.cc (address_operand): Permit vectors of addresses. * rtl.h (get_address_mode): Change return type to machine_mode. * rtlanal.cc (get_address_mode): Likewise. * simplify-rtx.cc (simplify_context::simplify_subreg): Handle vectors of addresses. * target.def (pointer_mode): Document vectors of addresses. (address_mode): Likewise. (valid_pointer_mode): Change scalar_int_mode to machine_mode. * targhooks.cc (default_addr_space_valid_pointer_mode): Change scalar_int_mode to machine_mode. * targhooks.h (default_addr_space_valid_pointer_mode): Likewise. --- gcc/doc/rtl.texi | 8 ++++++++ gcc/doc/tm.texi | 13 ++++++++++++- gcc/emit-rtl.cc | 30 ++++++++++++++++++++++++------ gcc/explow.cc | 12 ++++++++++-- gcc/recog.cc | 4 +++- gcc/rtl.h | 2 +- gcc/rtlanal.cc | 4 ++-- gcc/simplify-rtx.cc | 5 ++++- gcc/target.def | 13 ++++++++++--- gcc/targhooks.cc | 4 ++-- gcc/targhooks.h | 3 +-- 11 files changed, 77 insertions(+), 21 deletions(-) diff --git a/gcc/doc/rtl.texi b/gcc/doc/rtl.texi index f606b97389f..460d333e3da 100644 --- a/gcc/doc/rtl.texi +++ b/gcc/doc/rtl.texi @@ -2369,6 +2369,14 @@ a unit of memory is accessed. @var{alias} specifies an alias set for the reference. In general two items are in different alias sets if they cannot reference the same memory address. +The @var{addr} expression can represent either a scalar base address, or a +vector of addresses (using an appropriate vector mode). In the vector case, +the number of lanes in the address mode must match the number of lanes in the +data mode @var{m}. Each vector lane can be viewed as an independent, parallel +memory access of the corresponding scalar mode. If multiple lanes represent +the same, or overlapping, memory location then the effect is undefined. All +lanes are assumed to have the same attributes (address space, volatile, etc.). + The construct @code{(mem:BLK (scratch))} is considered to alias all other memories. Thus it may be used as a memory barrier in epilogue stack deallocation patterns. diff --git a/gcc/doc/tm.texi b/gcc/doc/tm.texi index fc9acdb38da..07e9265585c 100644 --- a/gcc/doc/tm.texi +++ b/gcc/doc/tm.texi @@ -11497,16 +11497,27 @@ c_register_addr_space ("__ea", ADDR_SPACE_EA); @deftypefn {Target Hook} scalar_int_mode TARGET_ADDR_SPACE_POINTER_MODE (addr_space_t @var{address_space}) Define this to return the machine mode to use for pointers to @var{address_space} if the target supports named address spaces. + +Targets that support vectors of pointers should return only the scalar base +mode here, but permit the vector modes in +@code{TARGET_ADDR_SPACE_VALID_POINTER_MODE}. + The default version of this hook returns @code{ptr_mode}. @end deftypefn @deftypefn {Target Hook} scalar_int_mode TARGET_ADDR_SPACE_ADDRESS_MODE (addr_space_t @var{address_space}) Define this to return the machine mode to use for addresses in @var{address_space} if the target supports named address spaces. + +Targets that support vectors of addresses should return only the scalar base +mode here, but permit the vector modes in +@code{TARGET_ADDR_SPACE_LEGITIMATE_ADDRESS_P} and +@code{TARGET_ADDR_SPACE_LEGITIMIZE_ADDRESS}. + The default version of this hook returns @code{Pmode}. @end deftypefn -@deftypefn {Target Hook} bool TARGET_ADDR_SPACE_VALID_POINTER_MODE (scalar_int_mode @var{mode}, addr_space_t @var{as}) +@deftypefn {Target Hook} bool TARGET_ADDR_SPACE_VALID_POINTER_MODE (machine_mode @var{mode}, addr_space_t @var{as}) Define this to return nonzero if the port can handle pointers with machine mode @var{mode} to address space @var{as}. This target hook is the same as the @code{TARGET_VALID_POINTER_MODE} target hook, diff --git a/gcc/emit-rtl.cc b/gcc/emit-rtl.cc index 4a23eaefe02..ea4d757bed3 100644 --- a/gcc/emit-rtl.cc +++ b/gcc/emit-rtl.cc @@ -540,14 +540,29 @@ gen_rtx_CONST_INT (machine_mode mode ATTRIBUTE_UNUSED, HOST_WIDE_INT arg) return *slot; } +/* Return a CONST_INT, CONST_POLY_INT, or CONST_VECTOR containing the given + value. In the case of a vector, the value will be duplicated. */ + rtx gen_int_mode (poly_int64 c, machine_mode mode) { - c = trunc_int_for_mode (c, mode); + machine_mode s_mode = (VECTOR_MODE_P (mode) ? GET_MODE_INNER (mode) : mode); + rtx val; + + c = trunc_int_for_mode (c, s_mode); if (c.is_constant ()) - return GEN_INT (c.coeffs[0]); - unsigned int prec = GET_MODE_PRECISION (as_a (mode)); - return immed_wide_int_const (poly_wide_int::from (c, prec, SIGNED), mode); + val = GEN_INT (c.coeffs[0]); + else + { + unsigned int prec = GET_MODE_PRECISION (as_a (s_mode)); + val = immed_wide_int_const (poly_wide_int::from (c, prec, SIGNED), + s_mode); + } + + if (VECTOR_MODE_P (mode)) + val = gen_const_vec_duplicate (mode, val); + + return val; } /* CONST_DOUBLEs might be created from pairs of integers, or from @@ -2381,7 +2396,7 @@ adjust_address_1 (rtx memref, machine_mode mode, poly_int64 offset, { rtx addr = XEXP (memref, 0); rtx new_rtx; - scalar_int_mode address_mode; + machine_mode address_mode; class mem_attrs attrs (*get_mem_attrs (memref)), *defattrs; unsigned HOST_WIDE_INT max_align; #ifdef POINTERS_EXTEND_UNSIGNED @@ -2415,7 +2430,10 @@ adjust_address_1 (rtx memref, machine_mode mode, poly_int64 offset, /* Convert a possibly large offset to a signed value within the range of the target address space. */ address_mode = get_address_mode (memref); - offset = trunc_int_for_mode (offset, address_mode); + if (!VECTOR_MODE_P (address_mode)) + offset = trunc_int_for_mode (offset, address_mode); + else + offset = trunc_int_for_mode (offset, GET_MODE_INNER (address_mode)); if (adjust_address) { diff --git a/gcc/explow.cc b/gcc/explow.cc index ec2bc9bc9b6..9c7c834aaab 100644 --- a/gcc/explow.cc +++ b/gcc/explow.cc @@ -298,8 +298,13 @@ convert_memory_address_addr_space_1 (scalar_int_mode to_mode ATTRIBUTE_UNUSED, bool in_const ATTRIBUTE_UNUSED, bool no_emit ATTRIBUTE_UNUSED) { + machine_mode xmode = GET_MODE (x); + bool vecaddr_p = VECTOR_MODE_P (GET_MODE (x)); + if (vecaddr_p) + xmode = GET_MODE_INNER (xmode); + #ifndef POINTERS_EXTEND_UNSIGNED - gcc_assert (GET_MODE (x) == to_mode || GET_MODE (x) == VOIDmode); + gcc_assert (xmode == to_mode || xmode == VOIDmode); return x; #else /* defined(POINTERS_EXTEND_UNSIGNED) */ scalar_int_mode pointer_mode, address_mode, from_mode; @@ -307,9 +312,12 @@ convert_memory_address_addr_space_1 (scalar_int_mode to_mode ATTRIBUTE_UNUSED, enum rtx_code code; /* If X already has the right mode, just return it. */ - if (GET_MODE (x) == to_mode) + if (xmode == to_mode) return x; + /* TODO: support vector conversions. */ + gcc_assert (!vecaddr_p); + pointer_mode = targetm.addr_space.pointer_mode (as); address_mode = targetm.addr_space.address_mode (as); from_mode = to_mode == pointer_mode ? address_mode : pointer_mode; diff --git a/gcc/recog.cc b/gcc/recog.cc index f9cff68e457..f0553fdfd05 100644 --- a/gcc/recog.cc +++ b/gcc/recog.cc @@ -1632,7 +1632,9 @@ address_operand (rtx op, machine_mode mode) { /* Wrong mode for an address expr. */ if (GET_MODE (op) != VOIDmode - && ! SCALAR_INT_MODE_P (GET_MODE (op))) + && !(SCALAR_INT_MODE_P (GET_MODE (op)) + || (VECTOR_MODE_P (GET_MODE (op)) + && SCALAR_INT_MODE_P (GET_MODE_INNER (GET_MODE (op)))))) return false; return memory_address_p (mode, op); diff --git a/gcc/rtl.h b/gcc/rtl.h index 01b4c6fd957..380e32b1fd4 100644 --- a/gcc/rtl.h +++ b/gcc/rtl.h @@ -3691,7 +3691,7 @@ inline rtx single_set (const rtx_insn *insn) return single_set_2 (insn, PATTERN (insn)); } -extern scalar_int_mode get_address_mode (rtx mem); +extern machine_mode get_address_mode (rtx mem); extern bool rtx_addr_can_trap_p (const_rtx); extern bool nonzero_address_p (const_rtx); extern bool rtx_unstable_p (const_rtx); diff --git a/gcc/rtlanal.cc b/gcc/rtlanal.cc index 5274a5c59cf..f741093f70e 100644 --- a/gcc/rtlanal.cc +++ b/gcc/rtlanal.cc @@ -6295,7 +6295,7 @@ low_bitmask_len (machine_mode mode, unsigned HOST_WIDE_INT m) /* Return the mode of MEM's address. */ -scalar_int_mode +machine_mode get_address_mode (rtx mem) { machine_mode mode; @@ -6303,7 +6303,7 @@ get_address_mode (rtx mem) gcc_assert (MEM_P (mem)); mode = GET_MODE (XEXP (mem, 0)); if (mode != VOIDmode) - return as_a (mode); + return mode; return targetm.addr_space.address_mode (MEM_ADDR_SPACE (mem)); } diff --git a/gcc/simplify-rtx.cc b/gcc/simplify-rtx.cc index 92a2a6e954a..c62eba41743 100644 --- a/gcc/simplify-rtx.cc +++ b/gcc/simplify-rtx.cc @@ -8679,7 +8679,10 @@ simplify_context::simplify_subreg (machine_mode outermode, rtx op, && (! MEM_VOLATILE_P (op) || ! have_insn_for (SET, innermode)) && !(STRICT_ALIGNMENT && MEM_ALIGN (op) < GET_MODE_ALIGNMENT (outermode)) - && known_le (outersize, innersize)) + && known_le (outersize, innersize) + /* You can't get the Nth element of a vector of addresses by adding + offsets. */ + && !VECTOR_MODE_P (GET_MODE (XEXP (op, 0)))) return adjust_address_nv (op, outermode, byte); /* Handle complex or vector values represented as CONCAT or VEC_CONCAT diff --git a/gcc/target.def b/gcc/target.def index 884fe1bd57e..283b81220d1 100644 --- a/gcc/target.def +++ b/gcc/target.def @@ -3448,7 +3448,10 @@ HOOK_VECTOR (TARGET_ADDR_SPACE_HOOKS, addr_space) DEFHOOK (pointer_mode, "Define this to return the machine mode to use for pointers to\n\ -@var{address_space} if the target supports named address spaces.\n\ +@var{address_space} if the target supports named address spaces.\n\n\ +Targets that support vectors of pointers should return only the scalar base\n\ +mode here, but permit the vector modes in\n\ +@code{TARGET_ADDR_SPACE_VALID_POINTER_MODE}.\n\n\ The default version of this hook returns @code{ptr_mode}.", scalar_int_mode, (addr_space_t address_space), default_addr_space_pointer_mode) @@ -3457,7 +3460,11 @@ The default version of this hook returns @code{ptr_mode}.", DEFHOOK (address_mode, "Define this to return the machine mode to use for addresses in\n\ -@var{address_space} if the target supports named address spaces.\n\ +@var{address_space} if the target supports named address spaces.\n\n\ +Targets that support vectors of addresses should return only the scalar base\n\ +mode here, but permit the vector modes in\n\ +@code{TARGET_ADDR_SPACE_LEGITIMATE_ADDRESS_P} and\n\ +@code{TARGET_ADDR_SPACE_LEGITIMIZE_ADDRESS}.\n\n\ The default version of this hook returns @code{Pmode}.", scalar_int_mode, (addr_space_t address_space), default_addr_space_address_mode) @@ -3473,7 +3480,7 @@ except that it includes explicit named address space support. The default\n\ version of this hook returns true for the modes returned by either the\n\ @code{TARGET_ADDR_SPACE_POINTER_MODE} or @code{TARGET_ADDR_SPACE_ADDRESS_MODE}\n\ target hooks for the given address space.", - bool, (scalar_int_mode mode, addr_space_t as), + bool, (machine_mode mode, addr_space_t as), default_addr_space_valid_pointer_mode) /* True if an address is a valid memory address to a given named address diff --git a/gcc/targhooks.cc b/gcc/targhooks.cc index 388f696c8aa..e89e9feb79e 100644 --- a/gcc/targhooks.cc +++ b/gcc/targhooks.cc @@ -1728,10 +1728,10 @@ default_addr_space_address_mode (addr_space_t addrspace ATTRIBUTE_UNUSED) To match the above, the same modes apply to all address spaces. */ bool -default_addr_space_valid_pointer_mode (scalar_int_mode mode, +default_addr_space_valid_pointer_mode (machine_mode mode, addr_space_t as ATTRIBUTE_UNUSED) { - return targetm.valid_pointer_mode (mode); + return targetm.valid_pointer_mode (as_a (mode)); } /* Some places still assume that all pointer or address modes are the diff --git a/gcc/targhooks.h b/gcc/targhooks.h index 86de6ef8a69..798a84d0e13 100644 --- a/gcc/targhooks.h +++ b/gcc/targhooks.h @@ -206,8 +206,7 @@ extern bool default_valid_pointer_mode (scalar_int_mode); extern bool default_ref_may_alias_errno (class ao_ref *); extern scalar_int_mode default_addr_space_pointer_mode (addr_space_t); extern scalar_int_mode default_addr_space_address_mode (addr_space_t); -extern bool default_addr_space_valid_pointer_mode (scalar_int_mode, - addr_space_t); +extern bool default_addr_space_valid_pointer_mode (machine_mode, addr_space_t); extern bool default_addr_space_legitimate_address_p (machine_mode, rtx, bool, addr_space_t, code_helper); extern rtx default_addr_space_legitimize_address (rtx, rtx, machine_mode, From patchwork Wed Jul 8 14:46:32 2026 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Andrew Stubbs X-Patchwork-Id: 138756 Return-Path: X-Original-To: patchwork@sourceware.org Delivered-To: patchwork@sourceware.org Received: from vm01.sourceware.org (localhost [IPv6:::1]) by sourceware.org (Postfix) with ESMTP id F2EAF4BA2E19 for ; Wed, 8 Jul 2026 14:47:55 +0000 (GMT) DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org F2EAF4BA2E19 Authentication-Results: sourceware.org; dkim=pass (2048-bit key, secure) header.d=baylibre.com header.i=@baylibre.com header.a=rsa-sha256 header.s=google header.b=O7QX2QIP X-Original-To: gcc-patches@gcc.gnu.org Delivered-To: gcc-patches@gcc.gnu.org Received: from mail-wr1-x42d.google.com (mail-wr1-x42d.google.com [IPv6:2a00:1450:4864:20::42d]) by sourceware.org (Postfix) with ESMTPS id 3AAE54BA2E09 for ; Wed, 8 Jul 2026 14:46:49 +0000 (GMT) DMARC-Filter: OpenDMARC Filter v1.4.2 sourceware.org 3AAE54BA2E09 Authentication-Results: sourceware.org; dmarc=none (p=none dis=none) header.from=baylibre.com Authentication-Results: sourceware.org; spf=pass smtp.mailfrom=baylibre.com ARC-Filter: OpenARC Filter v1.0.0 sourceware.org 3AAE54BA2E09 Authentication-Results: sourceware.org; arc=none smtp.remote-ip=2a00:1450:4864:20::42d ARC-Seal: i=1; a=rsa-sha256; d=sourceware.org; s=key; t=1783522009; cv=none; b=vSSy/bR8i6tZYpwsvCWno6ALdUqudJVnBOxlt+bVW4E7a5bNt6swjiz7y+VB+sCMP2cU7CDHCwXzaHX5RoFarnT2hDQ1nRs14+b+41hDDGq+3b1zgc3jMe6Ht6qTyB8sa0OIyw9JNNlQHQqFH4LhOo8W5GmKK0Vg90GIxnZGKb8= ARC-Message-Signature: i=1; a=rsa-sha256; d=sourceware.org; s=key; t=1783522009; c=relaxed/simple; bh=YMLq6So6Mnh5/7rFUpQF1KsItfon4XSMufl92cCPY4Q=; h=DKIM-Signature:From:To:Subject:Date:Message-ID:MIME-Version; b=vE5+ylQS7VHKroXqtf5FXh8VYAJDPvAQb/V8gJH3tcOgi1hyi5Fdh5DXHy1ReOHq/lfHudwxmoVT71bSPuMdbTjvtn6eC7psHB+aCDw6F3ICv0zMsfeg3GXpBz3nkEtRAGr8Y1wM+fWe65M3p26ZAplv13D71qq2l4lp3K2D40Y= ARC-Authentication-Results: i=1; sourceware.org; dkim=pass (2048-bit key, secure) header.d=baylibre.com header.i=@baylibre.com header.a=rsa-sha256 header.s=google header.b=O7QX2QIP DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org 3AAE54BA2E09 Received: by mail-wr1-x42d.google.com with SMTP id ffacd0b85a97d-4720f3bf164so1010110f8f.1 for ; Wed, 08 Jul 2026 07:46:49 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=baylibre.com; s=google; t=1783522008; x=1784126808; darn=gcc.gnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:to:from:from:to:cc:subject:date:message-id :reply-to:content-type; bh=7RJfxP4f/SgFa+leHsCMFdBU/5IeO5NwKBLHhfZVyPk=; b=O7QX2QIPXzGN37iVgrlfDXG9zCM9XgDa5exUXqzhtZfRn66fqeygLmQIGbW65yk6qz 14/d27uxjVEk5RpuuxoedTwrEVbqjyGD618e1grMz/eaGztxrGgxuOPz3HVJYcLbImNB hd1v/Qx8uIxYDxEhoujZ0eZnzFYKPJagPf2Q+i99oC0fxFiZFKAeHvmIQgtYhAcnPIlR xqiYXyx1qwCA9tgghGV7IfRtDVJTWPkm1CXpTZHCmiK6a6osR36MY316FmvLVgY55LCW jGSMV/CqXfP3MNVU7kT6+JgrU0EUT0rKLQIOQbZWwR5TQ2nHdka+4IwtgzmFVLrv9HKO hobQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783522008; x=1784126808; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:to:from:x-gm-gg:x-gm-message-state:from:to :cc:subject:date:message-id:reply-to:content-type; bh=7RJfxP4f/SgFa+leHsCMFdBU/5IeO5NwKBLHhfZVyPk=; b=qTIfOvHOYirBz0WcOgSxaGg0fZ6SjmI/m3S5daryk57uqAVYGW375yTebeh0qv7grO 8iWWq55xAsfXfdA8ASOlqM6vwBcA3IcQ9aByVTLIZY4i9xHFbykm8YLXu5CI0/JYuAPm 4UziJnkHDtD0iRBW8CvFJYHO7TCUQ45twiQwdeYbi1R9EigqBxnIOiEx19IgWTw0ODVW PVdiNdU7TzFZE6vwPrG5Z/l5MBkYiwvEOTOO09+dQL6l0nncjfORqWSbC+vy/POXfTAI DKRWkJvcbhVQHu8bcQtlMTkgwULXyS+sl59LzH3r9YMRi9L1Z/e+PpOEjW5FhX4/HISA yt8g== X-Gm-Message-State: AOJu0YxqtRzilz3zDhz5QpdyZOhkpbcHhKYadOs/HuKYFCNbMhnp4PBs 0J+6Px5VgzTCDQ+s2+z7T+FsuOxbP/yjWmwhOPL60CmDO19ZNWzQtY+x/FtzLJEXydjCd7vHeZX yD3Zd X-Gm-Gg: AfdE7ckwWQPBm7xco618UJCEuVvVtg+8XUwApArNIkUASE0xZgYs0udSZryKOJDLaz4 upjdkogIfBq+Y/ehXOfhJYaARa0qydNo4cEkAJKJNOlT3yVew7BZ2aA59jAqy5+BUoKdUnThBXq tu+Uekjpa/z/Vz+AAf0lMaTyWyhevUPne3956JIwjjvWEZTzWsJ7BNj305ptO+pnuCxNPENuih0 28PIQMl4lRngvkgK7LsgvD42uLtThbQjDtRjtgTacxvehXkLJjU4Whn1muahqhc4e52/8pcyQfe HWDEmWkWF3msbuMbopn5gKNLdCupJeB/HewqKHy63p6eey1+kkTYcEt29HochfFOvgJfetSxdds E0xXNAQaAx6fro+WkYv/uPi+3e/KX0VG8fh1JFvlHpsnPQEmGDN5lvJqZLYp2l+eiGBY9r/tNYg GUxJvfzuPYWHupw7g= X-Received: by 2002:a5d:5885:0:b0:46f:8561:fe60 with SMTP id ffacd0b85a97d-47de9a0e419mr8802672f8f.17.1783522006815; Wed, 08 Jul 2026 07:46:46 -0700 (PDT) Received: from vbuild-02.baylibre ([217.13.61.132]) by smtp.googlemail.com with ESMTPSA id ffacd0b85a97d-47a9e4d83bdsm43362109f8f.13.2026.07.08.07.46.45 for (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 08 Jul 2026 07:46:46 -0700 (PDT) From: Andrew Stubbs To: gcc-patches@gcc.gnu.org Subject: [PATCH 2/3] amdgcn: Implement "(mem (reg:))" Date: Wed, 8 Jul 2026 14:46:32 +0000 Message-ID: <20260708144633.1530935-3-ams@baylibre.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260708144633.1530935-1-ams@baylibre.com> References: <20260708144633.1530935-1-ams@baylibre.com> MIME-Version: 1.0 X-Spam-Status: No, score=-11.3 required=5.0 tests=BAYES_00, DKIM_SIGNED, DKIM_VALID, DKIM_VALID_AU, DKIM_VALID_EF, GIT_PATCH_0, RCVD_IN_DNSWL_NONE, SCC_5_SHORT_WORD_LINES, SPF_HELO_NONE, SPF_PASS, TXREP shortcircuit=no autolearn=ham autolearn_force=no version=3.4.6 X-Spam-Checker-Version: SpamAssassin 3.4.6 (2021-04-09) on sourceware.org X-BeenThere: gcc-patches@gcc.gnu.org X-Mailman-Version: 2.1.30 Precedence: list List-Id: Gcc-patches mailing list List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: gcc-patches-bounces~patchwork=sourceware.org@gcc.gnu.org This patch modifies the amdgcn MEM handling to allow, recognise, and output instructions that use vectors of addresses. Prior to this, amdgcn was forced to use gather/scatter to load/store vectors even though that is not really the most natural for the ISA. The code for validating MEMs and address registers is completely refactored. The "mov" instructions now include vector memory loads and stores for both contiguous and non-contiguous accesses. The gather/scatter patterns are removed and the expanders rewritten to use the general "mov" instructions. As a result, unconditional loads and stores are now represented with MEM in a much more conventional way, instead of being deconstructed and hidden inside UNSPECs. Sadly, there's still no way to indicate that some lanes of a MEM used in a masked load are invalid, so there's no way to stop LRA from "reloading" a MEM inside a vec_merge construct and putting it in a separate step (leading to memory faults). Therefore masked loads ("mov_exec") still have to use the UNSPEC trick ("load_exec") to prevent this happening. (The same issue should also affect other back-ends that have masked loads, but I believe AMD GCN is unique in having early-clobber constraints in such patterns, which are triggering the reloads, so those other targets are getting away with it.) gcc/ChangeLog: * config/gcn/constraints.md (RF): Limit to scalar addresses. (Rf): New. (Pf): New. (RL): Limit to scalar addresses. (Rl): New. (Pl): New. (RM): Limit to scalar addresses. (Rm): New. (Pm): New. (RA): New. * config/gcn/gcn-protos.h (gcn_gen_vector_mem): New prototype. (gcn_flat_address_p): Add "strict" parameter. (gcn_global_address_p): Add "strict" and mode parameter. (gcn_insn_base_reg_class): New prototype. (gcn_scalar_flat_address_p): Add "strict" parameter. * config/gcn/gcn-valu.md (scatter_store): Delete define_subst. (mov): Recognize address vectors. (*mov_1reg): Rework from the old "mov" patterns. (*mov_2reg): Likewise. (*mov_4reg): Likewise. (@mov_exec): New expander to separate masked loads from others. (*mov_1reg_exec): Likewise. (*mov_2reg_exec): Likewise. (*mov_4reg_exec): Likewise. (load_exec): New. (gather_load): Rework to use emit_move_insn. (gather_load): Likewise. (gather_expr): Delete. (gather_insn_1offset): Delete. (gather_insn_1offset_ds): Delete. (gather_insn_2offsets): Delete. (scatter_expr): Delete. (scatter_insn_1offset): Delete. (scatter_insn_1offset_ds): Delete. (scatter_insn_2offsets): Delete. (scatter_store): Rework to use emit_move_insn. (scatter_store): Likewise. (maskloaddi): Rework to use mov*_exec. (maskstoredi): Likewise. (mask_gather_load): Likewise. (mask_gather_load): Likewise. (mask_scatter_store): Likewise. (mask_scatter_store): Likewise. * config/gcn/gcn.cc (GEN_VNM): Delete gather_expr case. (gcn_address_register_p): Rename ... (gcn_scalar_address_register_p): ... to this. (gcn_vec_address_register_p): Support address vectors. (gcn_auto_address_register_p): New. (gcn_flat_address_p): Refactor code from gcn_addr_space_legitimate_address_p here, and support address vectors. (gcn_scalar_flat_address_p): Likewise. (gcn_ds_address_p): Likewise. (gcn_global_address_p): Likewise. (gcn_addr_space_legitimate_address_p): Likewise. (gcn_addr_space_valid_pointer_mode): New. (gcn_mode_code_base_reg_class): Refactor the code into ... (gcn_base_reg_class): ... here. (gcn_insn_base_reg_class): New. (gcn_expand_vector_init): Rework using mov*_exec. (gcn_expand_scalar_to_vector_address): Rework using vector MEM. (gcn_gen_vector_mem): New. (gcn_secondary_reload): Handle address vectors. (gcn_valid_move_p): Update calls to gcn_global_address_p. (move_callee_saved_registers): Rework using mov*_exec. (print_operand_address): Handle MEM changes. (print_operand): Likewise. (TARGET_ADDR_SPACE_VALID_POINTER_MODE): New. * config/gcn/gcn.h (INSN_BASE_REG_CLASS): New. * config/gcn/gcn.md (UNSPEC_MASKLOAD): New. (UNSPEC_GATHER): Delete. (UNSPEC_SCATTER): Delete. (atomic_fetch_): Add missing offsets. (atomic_): Likewise. (sync_compare_and_swap_insn): Likewise. --- gcc/config/gcn/constraints.md | 49 ++- gcc/config/gcn/gcn-protos.h | 10 +- gcc/config/gcn/gcn-valu.md | 745 ++++++++++++---------------------- gcc/config/gcn/gcn.cc | 529 +++++++++++++----------- gcc/config/gcn/gcn.h | 1 + gcc/config/gcn/gcn.md | 9 +- 6 files changed, 622 insertions(+), 721 deletions(-) diff --git a/gcc/config/gcn/constraints.md b/gcc/config/gcn/constraints.md index 4c3317a821d..e2fa688b373 100644 --- a/gcc/config/gcn/constraints.md +++ b/gcc/config/gcn/constraints.md @@ -116,7 +116,20 @@ (define_special_memory_constraint "RF" "Buffer memory address to flat memory." (and (match_code "mem") (match_test "AS_FLAT_P (MEM_ADDR_SPACE (op)) - && gcn_flat_address_p (XEXP (op, 0), mode)"))) + && gcn_flat_address_p (XEXP (op, 0), mode) + && !VECTOR_MODE_P (GET_MODE (XEXP (op, 0)))"))) + +(define_special_memory_constraint "Rf" + "Buffer memory address to flat memory." + (and (match_code "mem") + (match_test "AS_FLAT_P (MEM_ADDR_SPACE (op)) + && gcn_flat_address_p (XEXP (op, 0), mode) + && VECTOR_MODE_P (GET_MODE (XEXP (op, 0)))"))) + +(define_constraint "Pf" + "Integer matching ADDR_SPACE_FLAT" + (and (match_code "const_int") + (match_test "AS_FLAT_P (ival)"))) (define_special_memory_constraint "RS" "Buffer memory address to scalar flat memory." @@ -127,7 +140,19 @@ (define_special_memory_constraint "RS" (define_special_memory_constraint "RL" "Buffer memory address to LDS memory." (and (match_code "mem") - (match_test "AS_LDS_P (MEM_ADDR_SPACE (op))"))) + (match_test "AS_LDS_P (MEM_ADDR_SPACE (op)) + && !VECTOR_MODE_P (GET_MODE (XEXP (op, 0)))"))) + +(define_special_memory_constraint "Rl" + "Buffer memory address to LDS memory." + (and (match_code "mem") + (match_test "AS_LDS_P (MEM_ADDR_SPACE (op)) + && VECTOR_MODE_P (GET_MODE (XEXP (op, 0)))"))) + +(define_constraint "Pl" + "Integer matching ADDR_SPACE_LDS" + (and (match_code "const_int") + (match_test "AS_LDS_P (ival)"))) (define_special_memory_constraint "RG" "Buffer memory address to GDS memory." @@ -144,4 +169,22 @@ (define_special_memory_constraint "RM" "Memory address to global (main) memory." (and (match_code "mem") (match_test "AS_GLOBAL_P (MEM_ADDR_SPACE (op)) - && gcn_global_address_p (XEXP (op, 0))"))) + && gcn_global_address_p (XEXP (op, 0), GET_MODE (op)) + && !VECTOR_MODE_P (GET_MODE (XEXP (op, 0)))"))) + +(define_special_memory_constraint "Rm" + "Memory address to global (main) memory." + (and (match_code "mem") + (match_test "AS_GLOBAL_P (MEM_ADDR_SPACE (op)) + && gcn_global_address_p (XEXP (op, 0), GET_MODE (op)) + && VECTOR_MODE_P (GET_MODE (XEXP (op, 0)))"))) + +(define_constraint "Pm" + "Integer matching ADDR_SPACE_GLOBAL" + (and (match_code "const_int") + (match_test "AS_GLOBAL_P (ival)"))) + +(define_special_memory_constraint "RA" + "All memory types that use a scalar address." + (and (match_code "mem") + (match_test "!VECTOR_MODE_P (GET_MODE (XEXP (op, 0)))"))) diff --git a/gcc/config/gcn/gcn-protos.h b/gcc/config/gcn/gcn-protos.h index a4f5b58032e..e62bd01fb01 100644 --- a/gcc/config/gcn/gcn-protos.h +++ b/gcc/config/gcn/gcn-protos.h @@ -37,14 +37,17 @@ extern char * gcn_expand_dpp_distribute_odd_insn (machine_mode, const char *, extern void gcn_expand_epilogue (); extern rtx gcn_expand_scaled_offsets (addr_space_t as, rtx base, rtx offsets, rtx scale, bool unsigned_p, rtx exec); +extern rtx gcn_gen_vector_mem (machine_mode mode, addr_space_t as, + rtx scalarbase, rtx vectoroffsets, rtx scale, + bool unsigned_p, bool volatile_p, rtx exec); extern void gcn_expand_prologue (); extern rtx gcn_expand_reduc_scalar (machine_mode, rtx, int); extern rtx gcn_expand_scalar_to_vector_address (machine_mode, rtx, rtx, rtx); extern void gcn_expand_vector_init (rtx, rtx); -extern bool gcn_flat_address_p (rtx, machine_mode); +extern bool gcn_flat_address_p (rtx, machine_mode, bool strict = false); extern bool gcn_fp_constant_p (rtx, bool); extern rtx gcn_gen_undef (machine_mode); -extern bool gcn_global_address_p (rtx); +extern bool gcn_global_address_p (rtx, machine_mode, bool strict = false); extern tree gcn_goacc_adjust_private_decl (location_t, tree var, int level); extern tree gcn_goacc_create_worker_broadcast_record (tree record_type, bool sender, @@ -67,6 +70,7 @@ extern bool gcn_inline_constant_p (rtx); extern int gcn_inline_fp_constant_p (rtx, bool); extern reg_class gcn_mode_code_base_reg_class (machine_mode, addr_space_t, int, int); +extern reg_class gcn_insn_base_reg_class (rtx_insn *); extern rtx gcn_oacc_dim_pos (int dim); extern rtx gcn_oacc_dim_size (int dim); extern rtx gcn_operand_doublepart (machine_mode, rtx, int); @@ -74,7 +78,7 @@ extern rtx gcn_operand_part (machine_mode, rtx, int); extern bool gcn_regno_mode_code_ok_for_base_p (int, machine_mode, addr_space_t, int, int); extern reg_class gcn_regno_reg_class (int regno); -extern bool gcn_scalar_flat_address_p (rtx); +extern bool gcn_scalar_flat_address_p (rtx, bool strict = false); extern bool gcn_scalar_flat_mem_p (rtx); extern bool gcn_sgpr_move_p (rtx, rtx); extern bool gcn_stepped_zero_int_parallel_p (rtx op, int step); diff --git a/gcc/config/gcn/gcn-valu.md b/gcc/config/gcn/gcn-valu.md index edd5bf1f1d2..c63f768c791 100644 --- a/gcc/config/gcn/gcn-valu.md +++ b/gcc/config/gcn/gcn-valu.md @@ -304,8 +304,6 @@ (define_subst_attr "exec_clobber" "vec_merge_with_clobber" "" "_exec") (define_subst_attr "exec_vcc" "vec_merge_with_vcc" "" "_exec") -(define_subst_attr "exec_scatter" "scatter_store" - "" "_exec") (define_subst "vec_merge" [(set (match_operand:V_MOV 0) @@ -345,24 +343,6 @@ (define_subst "vec_merge_with_vcc" (and:DI (match_dup 3) (reg:DI EXEC_REG)))])]) -(define_subst "scatter_store" - [(set (mem:BLK (scratch)) - (unspec:BLK - [(match_operand 0) - (match_operand 1) - (match_operand 2) - (match_operand 3)] - UNSPEC_SCATTER))] - "" - [(set (mem:BLK (scratch)) - (unspec:BLK - [(match_dup 0) - (match_dup 1) - (match_dup 2) - (match_dup 3) - (match_operand:DI 4 "gcn_exec_reg_operand" "e")] - UNSPEC_SCATTER))]) - ;; }}} ;; {{{ Vector moves @@ -406,30 +386,31 @@ (define_expand "mov" && (!SUBREG_P (operands[1]) || !MEM_P (SUBREG_REG (operands[1])))); - if (MEM_P (operands[0]) && !lra_in_progress && !reload_completed) + if ((MEM_P (operands[0]) + && VECTOR_MODE_P (GET_MODE (XEXP (operands[0], 0)))) + || (MEM_P (operands[1]) + && VECTOR_MODE_P (GET_MODE (XEXP (operands[1], 0))))) + /* Use default expand. */ + ; + else if (MEM_P (operands[0]) && !lra_in_progress && !reload_completed) { operands[1] = force_reg (mode, operands[1]); rtx scratch = gen_rtx_SCRATCH (mode); - rtx a = gen_rtx_CONST_INT (VOIDmode, MEM_ADDR_SPACE (operands[0])); - rtx v = gen_rtx_CONST_INT (VOIDmode, MEM_VOLATILE_P (operands[0])); - rtx expr = gcn_expand_scalar_to_vector_address (mode, NULL, - operands[0], - scratch); - emit_insn (gen_scatter_expr (expr, operands[1], a, v)); - DONE; + operands[0] = gcn_expand_scalar_to_vector_address (mode, NULL, + operands[0], + scratch); } else if (MEM_P (operands[1]) && !lra_in_progress && !reload_completed) { rtx scratch = gen_rtx_SCRATCH (mode); - rtx a = gen_rtx_CONST_INT (VOIDmode, MEM_ADDR_SPACE (operands[1])); - rtx v = gen_rtx_CONST_INT (VOIDmode, MEM_VOLATILE_P (operands[1])); - rtx expr = gcn_expand_scalar_to_vector_address (mode, NULL, - operands[1], - scratch); - emit_insn (gen_gather_expr (operands[0], expr, a, v)); - DONE; + operands[1] = gcn_expand_scalar_to_vector_address (mode, NULL, + operands[1], + scratch); } - else if ((MEM_P (operands[0]) || MEM_P (operands[1]))) + else if ((MEM_P (operands[0]) + && !VECTOR_MODE_P (GET_MODE (XEXP (operands[0], 0)))) + || (MEM_P (operands[1]) + && !VECTOR_MODE_P (GET_MODE (XEXP (operands[1], 0))))) { gcc_assert (!reload_completed); rtx scratch = gen_reg_rtx (mode); @@ -448,88 +429,205 @@ (define_insn "mov_unspec" [(set_attr "type" "unknown") (set_attr "length" "0")]) -(define_insn "*mov" +(define_insn "*mov_1reg" [(set (match_operand:V_1REG 0 "nonimmediate_operand") (match_operand:V_1REG 1 "general_operand"))] + "!MEM_P (operands[0]) || REG_P (operands[1])" + {@ [cons: =0, 1; attrs: type, length, cdna, xnack] + [v ,vA;vop1 ,4 ,* ,* ] v_mov_b32\t%0, %1 + [v ,B ;vop1 ,8 ,* ,* ] ^ + [v ,Rf;flat ,12,* ,off] flat_load%o1\t%0, %A1%O1\;s_waitcnt\t0 + [&v ,Rf;flat ,12,* ,on ] ^ + [^a ,Rf;flat ,12,cdna2,off] ^ + [&^a,Rf;flat ,12,cdna2,on ] ^ + [Rf ,v ;flat ,12,* ,* ] flat_store%s0\t%A0, %1%O0 + [Rf ,a ;flat ,12,cdna2,* ] ^ + [v ,Rm;flat ,12,* ,off] global_load%o1\t%0, %A1%O1\;s_waitcnt\tvmcnt(0) + [&v ,Rm;flat ,12,* ,on ] ^ + [^a ,Rm;flat ,12,cdna2,off] ^ + [&^a,Rm;flat ,12,cdna2,on ] ^ + [Rm ,v ;flat ,12,* ,* ] global_store%s0\t%A0, %1%O0 + [Rm ,a ;flat ,12,cdna2,* ] ^ + [v ,Rl;ds ,12,* ,* ] ds_read%b1\t%0, %A1%O1\;s_waitcnt\tlgkmcnt(0) + [Rl ,v ;ds ,12,* ,* ] ds_write%b0\t%A0, %1%O0\;s_waitcnt\tlgkmcnt(0) + [v ,a ;vop3p_mai,8 ,* ,* ] v_accvgpr_read_b32\t%0, %1 + [$a ,v ;vop3p_mai,8 ,* ,* ] v_accvgpr_write_b32\t%0, %1 + [a ,a ;vop1 ,4 ,cdna2,* ] v_accvgpr_mov_b32\t%0, %1 + }) + +(define_insn "*mov_2reg" + [(set (match_operand:V_2REG 0 "nonimmediate_operand") + (match_operand:V_2REG 1 "general_operand"))] + "!MEM_P (operands[0]) || REG_P (operands[1])" + {@ [cons: =0, 1; attrs: type, length, cdna, xnack] + [v ,vDB;vmult,16,* ,* ] v_mov_b32\t%L0, %L1\;v_mov_b32\t%H0, %H1 + [v ,a ;vmult,16,* ,* ] v_accvgpr_read_b32\t%L0, %L1\;v_accvgpr_read_b32\t%H0, %H1 + [$a ,v ;vmult,16,* ,* ] v_accvgpr_write_b32\t%L0, %L1\;v_accvgpr_write_b32\t%H0, %H1 + [a ,a ;vmult,8 ,cdna2,* ] v_accvgpr_mov_b32\t%L0, %L1\;v_accvgpr_mov_b32\t%H0, %H1 + [v ,Rf ;flat ,12,* ,off] flat_load_dwordx2\t%0, %A1%O1\;s_waitcnt\t0 + [&v ,Rf ;flat ,12,* ,on ] ^ + [^a ,Rf ;flat ,12,cdna2,off] ^ + [&^a,Rf ;flat ,12,cdna2,on ] ^ + [Rf ,v ;flat ,12,* ,* ] flat_store_dwordx2\t%A0, %1%O0 + [Rf ,a ;flat ,12,cdna2,* ] ^ + [v ,Rm ;flat ,12,* ,off] global_load_dwordx2\t%0, %A1%O1\;s_waitcnt\tvmcnt(0) + [&v ,Rm ;flat ,12,* ,on ] ^ + [^a ,Rm ;flat ,12,cdna2,off] ^ + [&^a,Rm ;flat ,12,cdna2,on ] ^ + [Rm ,v ;flat ,12,* ,* ] global_store_dwordx2\t%A0, %1%O0 + [Rm ,a ;flat ,12,cdna2,* ] ^ + [v ,Rl ;ds ,12,* ,* ] ds_read_b64\t%0, %A1%O1\;s_waitcnt\tlgkmcnt(0) + [Rl ,v ;ds ,12,* ,* ] ds_write_b64\t%A0, %1%O0\;s_waitcnt\tlgkmcnt(0) + }) + +(define_insn "*mov_4reg" + [(set (match_operand:V_4REG 0 "nonimmediate_operand") + (match_operand:V_4REG 1 "general_operand"))] + "!MEM_P (operands[0]) || REG_P (operands[1])" + {@ [cons: =0, 1; attrs: type, length, cdna, xnack] + [v ,vDB;vmult,16,* ,* ] v_mov_b32\t%L0, %L1\; v_mov_b32\t%H0, %H1\; v_mov_b32\t%J0, %J1\; v_mov_b32\t%K0, %K1 + [v ,a ;vmult,32,* ,* ] v_accvgpr_read_b32\t%L0, %L1\; v_accvgpr_read_b32\t%H0, %H1\; v_accvgpr_read_b32\t%J0, %J1\; v_accvgpr_read_b32\t%K0, %K1 + [$a ,v ;vmult,32,* ,* ] v_accvgpr_write_b32\t%L0, %L1\;v_accvgpr_write_b32\t%H0, %H1\;v_accvgpr_write_b32\t%J0, %J1\;v_accvgpr_write_b32\t%K0, %K1 + [a ,a ;vmult,32,cdna2,* ] v_accvgpr_mov_b32\t%L0, %L1\; v_accvgpr_mov_b32\t%H0, %H1\; v_accvgpr_mov_b32\t%J0, %J1\; v_accvgpr_mov_b32\t%K0, %K1 + [v ,Rf ;flat ,12,* ,off] flat_load_dwordx4\t%0, %A1%O1\;s_waitcnt\t0 + [&v ,Rf ;flat ,12,* ,on ] ^ + [^a ,Rf ;flat ,12,cdna2,off] ^ + [&^a,Rf ;flat ,12,cdna2,on ] ^ + [Rf ,v ;flat ,12,* ,* ] flat_store_dwordx4\t%A0, %1%O0 + [Rf ,a ;flat ,12,cdna2,* ] ^ + [v ,Rm ;flat ,12,* ,off] global_load_dwordx4\t%0, %A1%O1\;s_waitcnt\tvmcnt(0) + [&v ,Rm ;flat ,12,* ,on ] ^ + [^a ,Rm ;flat ,12,cdna2,off] ^ + [&^a,Rm ;flat ,12,cdna2,on ] ^ + [Rm ,v ;flat ,12,* ,* ] global_store_dwordx4\t%A0, %1%O0 + [Rm ,a ;flat ,12,cdna2,* ] ^ + }) + +(define_expand "@mov_exec" + [(parallel [ + (set (match_operand:V_MOV 0 "nonimmediate_operand") + (vec_merge:V_MOV + (match_operand:V_MOV 1 "general_operand") + (match_operand:V_MOV 2 "general_or_unspec_operand") + (match_operand:DI 3 "gcn_exec_operand"))) + (clobber (match_scratch: 4))])] "" - {@ [cons: =0, 1; attrs: type, length, cdna] - [v ,vA;vop1 ,4,* ] v_mov_b32\t%0, %1 - [v ,B ;vop1 ,8,* ] ^ - [v ,a ;vop3p_mai,8,* ] v_accvgpr_read_b32\t%0, %1 - [$a ,v ;vop3p_mai,8,* ] v_accvgpr_write_b32\t%0, %1 - [a ,a ;vop1 ,4,cdna2] v_accvgpr_mov_b32\t%0, %1 + { + if (MEM_P (operands[1])) + { + /* Masked loads need an insn that doesn't use an explicit MEM. */ + rtx addr = XEXP (operands[1], 0); + rtx offset = const0_rtx; + if (GET_CODE (addr) == PLUS) + { + offset = XEXP (addr, 1); + addr = XEXP (addr, 0); + } + if (GET_CODE (offset) == VEC_DUPLICATE) + offset = XEXP (offset, 0); + if (CONST_VECTOR_P (offset)) + offset = CONST_VECTOR_ELT (offset, 0); + if (!REG_P (addr) || !CONST_INT_P (offset)) + { + addr = force_reg (mode, addr); + offset = const0_rtx; + } + emit_insn (gen_load_exec (operands[0], addr, offset, + GEN_INT (MEM_ADDR_SPACE (operands[1])), + operands[2], operands[3])); + DONE; + } }) -(define_insn "mov_exec" +(define_insn "*mov_1reg_exec" [(set (match_operand:V_1REG 0 "nonimmediate_operand") (vec_merge:V_1REG (match_operand:V_1REG 1 "general_operand") - (match_operand:V_1REG 2 "gcn_alu_or_unspec_operand") - (match_operand:DI 3 "register_operand"))) + (match_operand:V_1REG 2 "general_or_unspec_operand") + (match_operand:DI 3 "gcn_exec_operand"))) (clobber (match_scratch: 4))] "!MEM_P (operands[0]) || REG_P (operands[1])" - {@ [cons: =0, 1, 2, 3, =4; attrs: type, length] - [v,vA,U0,e ,X ;vop1 ,4 ] v_mov_b32\t%0, %1 - [v,B ,U0,e ,X ;vop1 ,8 ] v_mov_b32\t%0, %1 - [v,v ,vA,cV,X ;vop2 ,4 ] v_cndmask_b32\t%0, %2, %1, vcc - [v,vA,vA,Sv,X ;vop3a,8 ] v_cndmask_b32\t%0, %2, %1, %3 - [v,m ,U0,e ,&v;* ,16] # - [m,v ,U0,e ,&v;* ,16] # + {@ [cons: =0, 1, 2, 3, =4; attrs: type, length, cdna, xnack] + [v ,vA,U0,e ,X ;vop1 ,4 ,* ,* ] v_mov_b32\t%0, %1 + [v ,B ,U0,e ,X ;vop1 ,8 ,* ,* ] v_mov_b32\t%0, %1 + [Rf ,v ,U0,e ,X ;flat ,12,* ,* ] flat_store%s0\t%A0, %1%O0 + [Rf ,a ,U0,e ,X ;flat ,12,cdna2,* ] ^ + [Rm ,v ,U0,e ,X ;flat ,12,* ,* ] global_store%s0\t%A0, %1%O0 + [Rm ,a ,U0,e ,X ;flat ,12,cdna2,* ] ^ + [Rl ,v ,U0,e ,X ;ds ,12,* ,* ] ds_write%b0\t%A0, %1%O0\;s_waitcnt\tlgkmcnt(0) + [v ,v ,vA,cV,X ;vop2 ,4 ,* ,* ] v_cndmask_b32\t%0, %2, %1, vcc + [v ,vA,vA,Sv,X ;vop3a,8 ,* ,* ] v_cndmask_b32\t%0, %2, %1, %3 + [v ,RA,U0,e ,&v;* ,16,* ,* ] # + [RA ,v ,U0,e ,&v;* ,16,* ,* ] # }) -(define_insn "*mov" - [(set (match_operand:V_2REG 0 "nonimmediate_operand") - (match_operand:V_2REG 1 "general_operand"))] - "" - {@ [cons: =0, 1; attrs: length, cdna] - [v ,vDB;16,* ] v_mov_b32\t%L0, %L1\;v_mov_b32\t%H0, %H1 - [v ,a ;16,* ] v_accvgpr_read_b32\t%L0, %L1\;v_accvgpr_read_b32\t%H0, %H1 - [$a,v ;16,* ] v_accvgpr_write_b32\t%L0, %L1\;v_accvgpr_write_b32\t%H0, %H1 - [a ,a ;8 ,cdna2] v_accvgpr_mov_b32\t%L0, %L1\;v_accvgpr_mov_b32\t%H0, %H1 - } - [(set_attr "type" "vmult,vmult,vmult,vmult")]) - -(define_insn "mov_exec" +(define_insn "*mov_2reg_exec" [(set (match_operand:V_2REG 0 "nonimmediate_operand") (vec_merge:V_2REG (match_operand:V_2REG 1 "general_operand") - (match_operand:V_2REG 2 "gcn_alu_or_unspec_operand") - (match_operand:DI 3 "register_operand"))) + (match_operand:V_2REG 2 "general_or_unspec_operand") + (match_operand:DI 3 "gcn_exec_operand"))) (clobber (match_scratch: 4))] "!MEM_P (operands[0]) || REG_P (operands[1])" - {@ [cons: =0, 1, 2, 3, =4; attrs: type, length] - [v,vDB,U0 ,e ,X ;vmult,16] v_mov_b32\t%L0, %L1\;v_mov_b32\t%H0, %H1 - [v,v0 ,vDA0,cV,X ;vmult,16] v_cndmask_b32\t%L0, %L2, %L1, vcc\;v_cndmask_b32\t%H0, %H2, %H1, vcc - [v,v0 ,vDA0,Sv,X ;vmult,16] v_cndmask_b32\t%L0, %L2, %L1, %3\;v_cndmask_b32\t%H0, %H2, %H1, %3 - [v,m ,U0 ,e ,&v;* ,16] # - [m,v ,U0 ,e ,&v;* ,16] # - }) - -(define_insn "*mov_4reg" - [(set (match_operand:V_4REG 0 "nonimmediate_operand") - (match_operand:V_4REG 1 "general_operand"))] - "" - {@ [cons: =0, 1; attrs: type, length, cdna] - [v ,vDB;vmult,16,* ] v_mov_b32\t%L0, %L1\; v_mov_b32\t%H0, %H1\; v_mov_b32\t%J0, %J1\; v_mov_b32\t%K0, %K1 - [v ,a ;vmult,32,* ] v_accvgpr_read_b32\t%L0, %L1\; v_accvgpr_read_b32\t%H0, %H1\; v_accvgpr_read_b32\t%J0, %J1\; v_accvgpr_read_b32\t%K0, %K1 - [$a,v ;vmult,32,* ] v_accvgpr_write_b32\t%L0, %L1\;v_accvgpr_write_b32\t%H0, %H1\;v_accvgpr_write_b32\t%J0, %J1\;v_accvgpr_write_b32\t%K0, %K1 - [a ,a ;vmult,32,cdna2] v_accvgpr_mov_b32\t%L0, %L1\; v_accvgpr_mov_b32\t%H0, %H1\; v_accvgpr_mov_b32\t%J0, %J1\; v_accvgpr_mov_b32\t%K0, %K1 + {@ [cons: =0, 1, 2, 3, =4; attrs: type, length, cdna, xnack] + [v ,vDB,U0 ,e ,X ;vmult,16,* ,* ] v_mov_b32\t%L0, %L1\;v_mov_b32\t%H0, %H1 + [v ,v0 ,vDA0,cV,X ;vmult,16,* ,* ] v_cndmask_b32\t%L0, %L2, %L1, vcc\;v_cndmask_b32\t%H0, %H2, %H1, vcc + [v ,v0 ,vDA0,Sv,X ;vmult,16,* ,* ] v_cndmask_b32\t%L0, %L2, %L1, %3\;v_cndmask_b32\t%H0, %H2, %H1, %3 + [Rf ,v ,U0 ,e ,X ;flat ,12,* ,* ] flat_store_dwordx2\t%A0, %1%O0 + [Rf ,a ,U0 ,e ,X ;flat ,12,cdna2,* ] ^ + [Rm ,v ,U0 ,e ,X ;flat ,12,* ,* ] global_store_dwordx2\t%A0, %1%O0 + [Rm ,a ,U0 ,e ,X ;flat ,12,cdna2,* ] ^ + [Rl ,v ,U0 ,e ,X ;ds ,12,* ,* ] ds_write_b64\t%A0, %1%O0\;s_waitcnt\tlgkmcnt(0) + [v ,RA ,U0 ,e ,&v;* ,16,* ,* ] # + [RA ,v ,U0 ,e ,&v;* ,16,* ,* ] # }) -(define_insn "mov_exec" +(define_insn "*mov_4reg_exec" [(set (match_operand:V_4REG 0 "nonimmediate_operand") (vec_merge:V_4REG (match_operand:V_4REG 1 "general_operand") - (match_operand:V_4REG 2 "gcn_alu_or_unspec_operand") - (match_operand:DI 3 "register_operand"))) + (match_operand:V_4REG 2 "general_or_unspec_operand") + (match_operand:DI 3 "gcn_exec_operand"))) (clobber (match_scratch: 4))] "!MEM_P (operands[0]) || REG_P (operands[1])" - {@ [cons: =0, 1, 2, 3, =4; attrs: type, length] - [v,vDB,U0 ,e ,X ;vmult,32] v_mov_b32\t%L0, %L1\;v_mov_b32\t%H0, %H1\;v_mov_b32\t%J0, %J1\;v_mov_b32\t%K0, %K1 - [v,v0 ,vDA0,cV,X ;vmult,32] v_cndmask_b32\t%L0, %L2, %L1, vcc\;v_cndmask_b32\t%H0, %H2, %H1, vcc\;v_cndmask_b32\t%J0, %J2, %J1, vcc\;v_cndmask_b32\t%K0, %K2, %K1, vcc - [v,v0 ,vDA0,Sv,X ;vmult,32] v_cndmask_b32\t%L0, %L2, %L1, %3\;v_cndmask_b32\t%H0, %H2, %H1, %3\;v_cndmask_b32\t%J0, %J2, %J1, %3\;v_cndmask_b32\t%K0, %K2, %K1, %3 - [v,m ,U0 ,e ,&v;* ,32] # - [m,v ,U0 ,e ,&v;* ,32] # + {@ [cons: =0, 1, 2, 3, =4; attrs: type, length, cdna, xnack] + [v ,vDB,U0 ,e ,X ;vmult,32,* ,* ] v_mov_b32\t%L0, %L1\;v_mov_b32\t%H0, %H1\;v_mov_b32\t%J0, %J1\;v_mov_b32\t%K0, %K1 + [v ,v0 ,vDA0,cV,X ;vmult,32,* ,* ] v_cndmask_b32\t%L0, %L2, %L1, vcc\;v_cndmask_b32\t%H0, %H2, %H1, vcc\;v_cndmask_b32\t%J0, %J2, %J1, vcc\;v_cndmask_b32\t%K0, %K2, %K1, vcc + [v ,v0 ,vDA0,Sv,X ;vmult,32,* ,* ] v_cndmask_b32\t%L0, %L2, %L1, %3\;v_cndmask_b32\t%H0, %H2, %H1, %3\;v_cndmask_b32\t%J0, %J2, %J1, %3\;v_cndmask_b32\t%K0, %K2, %K1, %3 + [Rf ,v ,U0 ,e ,X ;flat ,12,* ,* ] flat_store_dwordx4\t%A0, %1%O0 + [Rf ,a ,U0 ,e ,X ;flat ,12,cdna2,* ] ^ + [Rm ,v ,U0 ,e ,X ;flat ,12,* ,* ] global_store_dwordx4\t%A0, %1%O0 + [Rm ,a ,U0 ,e ,X ;flat ,12,cdna2,* ] ^ + [v ,RA ,U0 ,e ,&v;* ,32,* ,* ] # + [RA ,v ,U0 ,e ,&v;* ,32,* ,* ] # + }) + +;; A mov_exec insn for partial loads that doesn't use an explicit MEM +;; and therefore can't get converted to an unmasked load by reload. +;; Other architectures don't seem to suffer from this problem, but that +;; may be because they don't have early-clobber constraints to cause reloads. +(define_insn "load_exec" + [(set (match_operand:V_MOV 0 "register_operand") + (vec_merge:V_MOV + (unspec:V_MOV [ + (plus: (match_operand: 1 "register_operand") + (match_operand:DI 2 "immediate_operand")) + (match_operand:SI 3 "immediate_operand") + (mem:BLK (const_int 0))] + UNSPEC_MASKLOAD) + (match_operand:V_MOV 4 "general_or_unspec_operand") + (match_operand:DI 5 "gcn_exec_operand")))] + "" + {@ [cons: =0, 1, 2, 3, 4, 5; attrs: type, length, cdna, xnack] + [v ,v,i,Pf,U0,e ;flat ,12,* ,off] flat_load%o0\t%0, %1 offset:%2\;s_waitcnt\t0 + [&v ,v,i,Pf,U0,e ;flat ,12,* ,on ] ^ + [^a ,v,i,Pf,U0,e ;flat ,12,cdna2,off] ^ + [&^a,v,i,Pf,U0,e ;flat ,12,cdna2,on ] ^ + [v ,v,i,Pm,U0,e ;flat ,12,* ,off] global_load%o0\t%0, %1, off offset:%2\;s_waitcnt\tvmcnt(0) + [&v ,v,i,Pm,U0,e ;flat ,12,* ,on ] ^ + [^a ,v,i,Pm,U0,e ;flat ,12,cdna2,off] ^ + [&^a,v,i,Pm,U0,e ;flat ,12,cdna2,on ] ^ + [v ,v,i,Pl,U0,e ;ds ,12,* ,* ] ds_read%b0\t%0, %1 offset:%2\;s_waitcnt\tlgkmcnt(0) }) ; A SGPR-base load looks like: @@ -588,7 +686,7 @@ (define_insn "@mov_sgprbase" [m,v ,&v;* ,12] # }) -; Expand scalar addresses into gather/scatter patterns +; Expand scalar base addresses into vectors of addresses (define_split [(set (match_operand:V_MOV 0 "memory_operand") @@ -596,16 +694,12 @@ (define_split [(match_operand:V_MOV 1 "general_operand")] UNSPEC_SGPRBASE)) (clobber (match_scratch: 2))] - "" - [(set (mem:BLK (scratch)) - (unspec:BLK [(match_dup 5) (match_dup 1) (match_dup 6) (match_dup 7)] - UNSPEC_SCATTER))] + "!VECTOR_MODE_P (GET_MODE (XEXP (operands[0], 0)))" + [(set (match_dup 0) (match_dup 1))] { - operands[5] = gcn_expand_scalar_to_vector_address (mode, NULL, + operands[0] = gcn_expand_scalar_to_vector_address (mode, NULL, operands[0], operands[2]); - operands[6] = gen_rtx_CONST_INT (VOIDmode, MEM_ADDR_SPACE (operands[0])); - operands[7] = gen_rtx_CONST_INT (VOIDmode, MEM_VOLATILE_P (operands[0])); }) (define_split @@ -615,18 +709,17 @@ (define_split (match_operand:V_MOV 2 "") (match_operand:DI 3 "gcn_exec_reg_operand"))) (clobber (match_scratch: 4))] - "" - [(set (mem:BLK (scratch)) - (unspec:BLK [(match_dup 5) (match_dup 1) - (match_dup 6) (match_dup 7) (match_dup 3)] - UNSPEC_SCATTER))] + "!VECTOR_MODE_P (GET_MODE (XEXP (operands[0], 0)))" + [(parallel [(set (match_dup 0) + (vec_merge:V_MOV (match_dup 1) (match_dup 2) (match_dup 3))) + (clobber (match_dup 4))])] { - operands[5] = gcn_expand_scalar_to_vector_address (mode, - operands[3], - operands[0], - operands[4]); - operands[6] = gen_rtx_CONST_INT (VOIDmode, MEM_ADDR_SPACE (operands[0])); - operands[7] = gen_rtx_CONST_INT (VOIDmode, MEM_VOLATILE_P (operands[0])); + rtx mem = gcn_expand_scalar_to_vector_address (mode, operands[3], + operands[0], operands[4]); + if (rtx_equal_p (operands[0], operands[2])) + operands[0] = operands[2] = mem; + else + operands[0] = mem; }) (define_split @@ -635,17 +728,12 @@ (define_split [(match_operand:V_MOV 1 "memory_operand")] UNSPEC_SGPRBASE)) (clobber (match_scratch: 2))] - "" - [(set (match_dup 0) - (unspec:V_MOV [(match_dup 5) (match_dup 6) (match_dup 7) - (mem:BLK (scratch))] - UNSPEC_GATHER))] + "!VECTOR_MODE_P (GET_MODE (XEXP (operands[1], 0)))" + [(set (match_dup 0) (match_dup 1))] { - operands[5] = gcn_expand_scalar_to_vector_address (mode, NULL, + operands[1] = gcn_expand_scalar_to_vector_address (mode, NULL, operands[1], operands[2]); - operands[6] = gen_rtx_CONST_INT (VOIDmode, MEM_ADDR_SPACE (operands[1])); - operands[7] = gen_rtx_CONST_INT (VOIDmode, MEM_VOLATILE_P (operands[1])); }) (define_split @@ -655,21 +743,14 @@ (define_split (match_operand:V_MOV 2 "") (match_operand:DI 3 "gcn_exec_reg_operand"))) (clobber (match_scratch: 4))] - "" - [(set (match_dup 0) - (vec_merge:V_MOV - (unspec:V_MOV [(match_dup 5) (match_dup 6) (match_dup 7) - (mem:BLK (scratch))] - UNSPEC_GATHER) - (match_dup 2) - (match_dup 3)))] + "!VECTOR_MODE_P (GET_MODE (XEXP (operands[1], 0)))" + [(const_int 0)] { - operands[5] = gcn_expand_scalar_to_vector_address (mode, - operands[3], - operands[1], - operands[4]); - operands[6] = gen_rtx_CONST_INT (VOIDmode, MEM_ADDR_SPACE (operands[1])); - operands[7] = gen_rtx_CONST_INT (VOIDmode, MEM_VOLATILE_P (operands[1])); + rtx mem = gcn_expand_scalar_to_vector_address (mode, operands[3], + operands[1], operands[4]); + emit_insn (gen_mov_exec (operands[0], mem, operands[2], + operands[3])); + DONE; }) ; TODO: Add zero/sign extending variants. @@ -976,38 +1057,6 @@ (define_expand "vec_init" ;; }}} ;; {{{ Scatter / Gather -;; GCN does not have an instruction for loading a vector from contiguous -;; memory so *all* loads and stores are eventually converted to scatter -;; or gather. -;; -;; GCC does not permit MEM to hold vectors of addresses, so we must use an -;; unspec. The unspec formats are as follows: -;; -;; (unspec:V?? -;; [(
) -;; () -;; () -;; (mem:BLK (scratch))] -;; UNSPEC_GATHER) -;; -;; (unspec:BLK -;; [(
) -;; () -;; () -;; () -;; ()] -;; UNSPEC_SCATTER) -;; -;; - Loads are expected to be wrapped in a vec_merge, so do not need . -;; - The mem:BLK does not contain any real information, but indicates that an -;; unknown memory read is taking place. Stores are expected to use a similar -;; mem:BLK outside the unspec. -;; - The address space and glc (volatile) fields are there to replace the -;; fields normally found in a MEM. -;; - Multiple forms of address expression are supported, below. -;; -;; TODO: implement combined gather and zero_extend, but only for -msram-ecc=on - (define_expand "gather_load" [(match_operand:V_MOV 0 "register_operand") (match_operand:DI 1 "register_operand") @@ -1016,17 +1065,10 @@ (define_expand "gather_load" (match_operand:SI 4 "gcn_alu_operand")] "" { - rtx addr = gcn_expand_scaled_offsets (DEFAULT_ADDR_SPACE, operands[1], - operands[2], operands[4], - INTVAL (operands[3]), NULL); - - if (GET_MODE (addr) == mode) - emit_insn (gen_gather_insn_1offset (operands[0], addr, const0_rtx, - const0_rtx, const0_rtx)); - else - emit_insn (gen_gather_insn_2offsets (operands[0], operands[1], - addr, const0_rtx, const0_rtx, - const0_rtx)); + emit_move_insn (operands[0], + gcn_gen_vector_mem (mode, DEFAULT_ADDR_SPACE, + operands[1], operands[2], operands[4], + INTVAL (operands[3]), false, NULL)); DONE; }) @@ -1038,121 +1080,13 @@ (define_expand "gather_load" (match_operand:SI 4 "gcn_alu_operand")] "" { - rtx addr = gcn_expand_scaled_offsets (DEFAULT_ADDR_SPACE, operands[1], - operands[2], operands[4], - INTVAL (operands[3]), NULL); - - emit_insn (gen_gather_insn_1offset (operands[0], addr, const0_rtx, - const0_rtx, const0_rtx)); + emit_move_insn (operands[0], + gcn_gen_vector_mem (mode, DEFAULT_ADDR_SPACE, + operands[1], operands[2], operands[4], + INTVAL (operands[3]), false, NULL)); DONE; }) -; Allow any address expression -(define_expand "gather_expr" - [(set (match_operand:V_MOV 0 "register_operand") - (unspec:V_MOV - [(match_operand 1 "") - (match_operand 2 "immediate_operand") - (match_operand 3 "immediate_operand") - (mem:BLK (scratch))] - UNSPEC_GATHER))] - "" - {}) - -(define_insn "gather_insn_1offset" - [(set (match_operand:V_MOV 0 "register_operand" "=v,a,&v,&a") - (unspec:V_MOV - [(plus: (match_operand: 1 "register_operand" " v,v, v, v") - (vec_duplicate: - (match_operand 2 "immediate_operand" " n,n, n, n"))) - (match_operand 3 "immediate_operand" " n,n, n, n") - (match_operand 4 "immediate_operand" " n,n, n, n") - (mem:BLK (scratch))] - UNSPEC_GATHER))] - "(AS_FLAT_P (INTVAL (operands[3])) - && ((unsigned HOST_WIDE_INT)INTVAL(operands[2]) < 0x1000)) - || (AS_GLOBAL_P (INTVAL (operands[3])) - && (((unsigned HOST_WIDE_INT)INTVAL(operands[2]) + 0x1000) < 0x2000))" - { - addr_space_t as = INTVAL (operands[3]); - const char *glc = INTVAL (operands[4]) ? TARGET_GLC_NAME : ""; - - static char buf[200]; - if (AS_FLAT_P (as)) - sprintf (buf, "flat_load%%o0\t%%0, %%1 offset:%%2%s\;s_waitcnt\t0", glc); - else if (AS_GLOBAL_P (as)) - sprintf (buf, "global_load%%o0\t%%0, %%1, off offset:%%2%s\;" - "s_waitcnt\tvmcnt(0)", glc); - else - gcc_unreachable (); - - return buf; - } - [(set_attr "type" "flat") - (set_attr "flatmemaccess" "load") - (set_attr "length" "12") - (set_attr "cdna" "*,cdna2,*,cdna2") - (set_attr "xnack" "off,off,on,on")]) - -(define_insn "gather_insn_1offset_ds" - [(set (match_operand:V_MOV 0 "register_operand" "=v,a") - (unspec:V_MOV - [(plus: (match_operand: 1 "register_operand" " v,v") - (vec_duplicate: - (match_operand 2 "immediate_operand" " n,n"))) - (match_operand 3 "immediate_operand" " n,n") - (match_operand 4 "immediate_operand" " n,n") - (mem:BLK (scratch))] - UNSPEC_GATHER))] - "(AS_ANY_DS_P (INTVAL (operands[3])) - && ((unsigned HOST_WIDE_INT)INTVAL(operands[2]) < 0x10000))" - { - addr_space_t as = INTVAL (operands[3]); - static char buf[200]; - sprintf (buf, "ds_read%%b0\t%%0, %%1 offset:%%2%s\;s_waitcnt\tlgkmcnt(0)", - (AS_GDS_P (as) ? " gds" : "")); - return buf; - } - [(set_attr "type" "ds") - (set_attr "length" "12") - (set_attr "cdna" "*,cdna2")]) - -(define_insn "gather_insn_2offsets" - [(set (match_operand:V_MOV 0 "register_operand" "=v,a,&v,&a") - (unspec:V_MOV - [(plus: - (plus: - (vec_duplicate: - (match_operand:DI 1 "register_operand" "Sv,Sv,Sv,Sv")) - (sign_extend: - (match_operand: 2 "register_operand" " v, v, v, v"))) - (vec_duplicate: (match_operand 3 "immediate_operand" - " n, n, n, n"))) - (match_operand 4 "immediate_operand" " n, n, n, n") - (match_operand 5 "immediate_operand" " n, n, n, n") - (mem:BLK (scratch))] - UNSPEC_GATHER))] - "(AS_GLOBAL_P (INTVAL (operands[4])) - && (((unsigned HOST_WIDE_INT)INTVAL(operands[3]) + 0x1000) < 0x2000))" - { - addr_space_t as = INTVAL (operands[4]); - const char *glc = INTVAL (operands[5]) ? TARGET_GLC_NAME : ""; - - static char buf[200]; - if (AS_GLOBAL_P (as)) - sprintf (buf, "global_load%%o0\t%%0, %%2, %%1 offset:%%3%s\;" - "s_waitcnt\tvmcnt(0)", glc); - else - gcc_unreachable (); - - return buf; - } - [(set_attr "type" "flat") - (set_attr "flatmemaccess" "load") - (set_attr "length" "12") - (set_attr "cdna" "*,cdna2,*,cdna2") - (set_attr "xnack" "off,off,on,on")]) - (define_expand "scatter_store" [(match_operand:DI 0 "register_operand") (match_operand: 1 "register_operand") @@ -1161,17 +1095,10 @@ (define_expand "scatter_store" (match_operand:V_MOV 4 "register_operand")] "" { - rtx addr = gcn_expand_scaled_offsets (DEFAULT_ADDR_SPACE, operands[0], - operands[1], operands[3], - INTVAL (operands[2]), NULL); - - if (GET_MODE (addr) == mode) - emit_insn (gen_scatter_insn_1offset (addr, const0_rtx, operands[4], - const0_rtx, const0_rtx)); - else - emit_insn (gen_scatter_insn_2offsets (operands[0], addr, - const0_rtx, operands[4], - const0_rtx, const0_rtx)); + emit_move_insn (gcn_gen_vector_mem (mode, DEFAULT_ADDR_SPACE, + operands[0], operands[1], operands[3], + INTVAL (operands[2]), false, NULL), + operands[4]); DONE; }) @@ -1183,117 +1110,13 @@ (define_expand "scatter_store" (match_operand:V_MOV 4 "register_operand")] "" { - rtx addr = gcn_expand_scaled_offsets (DEFAULT_ADDR_SPACE, operands[0], - operands[1], operands[3], - INTVAL (operands[2]), NULL); - - emit_insn (gen_scatter_insn_1offset (addr, const0_rtx, operands[4], - const0_rtx, const0_rtx)); + emit_move_insn (gcn_gen_vector_mem (mode, DEFAULT_ADDR_SPACE, + operands[0], operands[1], operands[3], + INTVAL (operands[2]), false, NULL), + operands[4]); DONE; }) -; Allow any address expression -(define_expand "scatter_expr" - [(set (mem:BLK (scratch)) - (unspec:BLK - [(match_operand: 0 "") - (match_operand:V_MOV 1 "register_operand") - (match_operand 2 "immediate_operand") - (match_operand 3 "immediate_operand")] - UNSPEC_SCATTER))] - "" - {}) - -(define_insn "scatter_insn_1offset" - [(set (mem:BLK (scratch)) - (unspec:BLK - [(plus: (match_operand: 0 "register_operand" "v,v") - (vec_duplicate: - (match_operand 1 "immediate_operand" "n,n"))) - (match_operand:V_MOV 2 "register_operand" "v,a") - (match_operand 3 "immediate_operand" "n,n") - (match_operand 4 "immediate_operand" "n,n")] - UNSPEC_SCATTER))] - "(AS_FLAT_P (INTVAL (operands[3])) - && (INTVAL(operands[1]) == 0 - || ((unsigned HOST_WIDE_INT)INTVAL(operands[1]) < 0x1000))) - || (AS_GLOBAL_P (INTVAL (operands[3])) - && (((unsigned HOST_WIDE_INT)INTVAL(operands[1]) + 0x1000) < 0x2000))" - { - addr_space_t as = INTVAL (operands[3]); - const char *glc = INTVAL (operands[4]) ? TARGET_GLC_NAME : ""; - - static char buf[200]; - if (AS_FLAT_P (as)) - sprintf (buf, "flat_store%%s2\t%%0, %%2 offset:%%1%s", glc); - else if (AS_GLOBAL_P (as)) - sprintf (buf, "global_store%%s2\t%%0, %%2, off offset:%%1%s", glc); - else - gcc_unreachable (); - - return buf; - } - [(set_attr "type" "flat") - (set_attr "flatmemaccess" "store") - (set_attr "length" "12") - (set_attr "cdna" "*,cdna2")]) - -(define_insn "scatter_insn_1offset_ds" - [(set (mem:BLK (scratch)) - (unspec:BLK - [(plus: (match_operand: 0 "register_operand" "v,v") - (vec_duplicate: - (match_operand 1 "immediate_operand" "n,n"))) - (match_operand:V_MOV 2 "register_operand" "v,a") - (match_operand 3 "immediate_operand" "n,n") - (match_operand 4 "immediate_operand" "n,n")] - UNSPEC_SCATTER))] - "(AS_ANY_DS_P (INTVAL (operands[3])) - && ((unsigned HOST_WIDE_INT)INTVAL(operands[1]) < 0x10000))" - { - addr_space_t as = INTVAL (operands[3]); - static char buf[200]; - sprintf (buf, "ds_write%%b2\t%%0, %%2 offset:%%1%s\;s_waitcnt\tlgkmcnt(0)", - (AS_GDS_P (as) ? " gds" : "")); - return buf; - } - [(set_attr "type" "ds") - (set_attr "length" "12") - (set_attr "cdna" "*,cdna2")]) - -(define_insn "scatter_insn_2offsets" - [(set (mem:BLK (scratch)) - (unspec:BLK - [(plus: - (plus: - (vec_duplicate: - (match_operand:DI 0 "register_operand" "Sv,Sv")) - (sign_extend: - (match_operand: 1 "register_operand" "v,v"))) - (vec_duplicate: (match_operand 2 "immediate_operand" "n,n"))) - (match_operand:V_MOV 3 "register_operand" "v,a") - (match_operand 4 "immediate_operand" "n,n") - (match_operand 5 "immediate_operand" "n,n")] - UNSPEC_SCATTER))] - "(AS_GLOBAL_P (INTVAL (operands[4])) - && (((unsigned HOST_WIDE_INT)INTVAL(operands[2]) + 0x1000) < 0x2000))" - { - addr_space_t as = INTVAL (operands[4]); - const char *glc = INTVAL (operands[5]) ? TARGET_GLC_NAME : ""; - - static char buf[200]; - if (AS_GLOBAL_P (as)) - sprintf (buf, "global_store%%s3\t%%1, %%3, %%0 offset:%%2%s", glc); - else - gcc_unreachable (); - - return buf; - } - [(set_attr "type" "flat") - (set_attr "flatmemaccess" "store") - (set_attr "length" "12") - (set_attr "cdna" "*,cdna2")]) - ;; }}} ;; {{{ Permutations @@ -4120,13 +3943,10 @@ (define_expand "maskloaddi" "" { rtx exec = force_reg (DImode, operands[2]); - rtx addr = gcn_expand_scalar_to_vector_address + rtx mem = gcn_expand_scalar_to_vector_address (mode, exec, operands[1], gen_rtx_SCRATCH (mode)); - rtx as = gen_rtx_CONST_INT (VOIDmode, MEM_ADDR_SPACE (operands[1])); - rtx v = gen_rtx_CONST_INT (VOIDmode, MEM_VOLATILE_P (operands[1])); - - emit_insn (gen_gather_expr_exec (operands[0], addr, as, v, - gcn_gen_undef (mode), exec)); + emit_insn (gen_mov_exec (operands[0], mem, + gcn_gen_undef (mode), exec)); DONE; }) @@ -4137,11 +3957,10 @@ (define_expand "maskstoredi" "" { rtx exec = force_reg (DImode, operands[2]); - rtx addr = gcn_expand_scalar_to_vector_address + rtx mem = gcn_expand_scalar_to_vector_address (mode, exec, operands[0], gen_rtx_SCRATCH (mode)); - rtx as = gen_rtx_CONST_INT (VOIDmode, MEM_ADDR_SPACE (operands[0])); - rtx v = gen_rtx_CONST_INT (VOIDmode, MEM_VOLATILE_P (operands[0])); - emit_insn (gen_scatter_expr_exec (addr, operands[1], as, v, exec)); + emit_insn (gen_mov_exec (mem, operands[1], + gcn_gen_undef (mode), exec)); DONE; }) @@ -4156,51 +3975,30 @@ (define_expand "mask_gather_load" "" { rtx exec = force_reg (DImode, operands[5]); - - rtx addr = gcn_expand_scaled_offsets (DEFAULT_ADDR_SPACE, operands[1], - operands[2], operands[4], - INTVAL (operands[3]), exec); - - if (GET_MODE (addr) == mode) - emit_insn (gen_gather_insn_1offset_exec (operands[0], addr, - const0_rtx, const0_rtx, - const0_rtx, - gcn_gen_undef - (mode), - exec)); - else - emit_insn (gen_gather_insn_2offsets_exec (operands[0], operands[1], - addr, const0_rtx, - const0_rtx, const0_rtx, - gcn_gen_undef - (mode), - exec)); + rtx srcmem = gcn_gen_vector_mem (mode, DEFAULT_ADDR_SPACE, + operands[1], operands[2], operands[4], + INTVAL (operands[3]), false, exec); + emit_insn (gen_mov_exec (operands[0], srcmem, + gcn_gen_undef (mode), exec)); DONE; }) (define_expand "mask_gather_load" - [(set:V_MOV (match_operand:V_MOV 0 "register_operand") - (unspec:V_MOV - [(match_operand:DI 1 "register_operand") - (match_operand: 2 "register_operand") - (match_operand 3 "immediate_operand") - (match_operand:SI 4 "gcn_alu_operand") - (match_operand:DI 5 "") - (match_operand:V_MOV 6 "maskload_else_operand")] - UNSPEC_GATHER))] + [(match_operand:V_MOV 0 "register_operand") + (match_operand:DI 1 "register_operand") + (match_operand: 2 "register_operand") + (match_operand 3 "immediate_operand") + (match_operand:SI 4 "gcn_alu_operand") + (match_operand:DI 5 "") + (match_operand:V_MOV 6 "maskload_else_operand")] "" { rtx exec = force_reg (DImode, operands[5]); - - rtx addr = gcn_expand_scaled_offsets (DEFAULT_ADDR_SPACE, operands[1], - operands[2], operands[4], - INTVAL (operands[3]), exec); - - emit_insn (gen_gather_insn_1offset_exec (operands[0], addr, - const0_rtx, const0_rtx, - const0_rtx, - gcn_gen_undef (mode), - exec)); + rtx srcmem = gcn_gen_vector_mem (mode, DEFAULT_ADDR_SPACE, + operands[1], operands[2], operands[4], + INTVAL (operands[3]), false, exec); + emit_insn (gen_mov_exec (operands[0], srcmem, + gcn_gen_undef (mode), exec)); DONE; }) @@ -4214,21 +4012,11 @@ (define_expand "mask_scatter_store" "" { rtx exec = force_reg (DImode, operands[5]); - - rtx addr = gcn_expand_scaled_offsets (DEFAULT_ADDR_SPACE, operands[0], - operands[1], operands[3], - INTVAL (operands[2]), exec); - - if (GET_MODE (addr) == mode) - emit_insn (gen_scatter_insn_1offset_exec (addr, const0_rtx, - operands[4], const0_rtx, - const0_rtx, - exec)); - else - emit_insn (gen_scatter_insn_2offsets_exec (operands[0], addr, - const0_rtx, operands[4], - const0_rtx, const0_rtx, - exec)); + rtx destmem = gcn_gen_vector_mem (mode, DEFAULT_ADDR_SPACE, + operands[0], operands[1], operands[3], + INTVAL (operands[2]), false, exec); + emit_insn (gen_mov_exec (destmem, operands[4], + gcn_gen_undef (mode), exec)); DONE; }) @@ -4242,14 +4030,11 @@ (define_expand "mask_scatter_store" "" { rtx exec = force_reg (DImode, operands[5]); - - rtx addr = gcn_expand_scaled_offsets (DEFAULT_ADDR_SPACE, operands[0], - operands[1], operands[3], - INTVAL (operands[2]), exec); - - emit_insn (gen_scatter_insn_1offset_exec (addr, const0_rtx, - operands[4], const0_rtx, - const0_rtx, exec)); + rtx destmem = gcn_gen_vector_mem (mode, DEFAULT_ADDR_SPACE, + operands[0], operands[1], operands[3], + INTVAL (operands[2]), false, exec); + emit_insn (gen_mov_exec (destmem, operands[4], + gcn_gen_undef (mode), exec)); DONE; }) diff --git a/gcc/config/gcn/gcn.cc b/gcc/config/gcn/gcn.cc index ed27acfb8a1..c110049bbbc 100644 --- a/gcc/config/gcn/gcn.cc +++ b/gcc/config/gcn/gcn.cc @@ -1431,8 +1431,6 @@ GEN_VN (addc,si3, A(rtx dest, rtx src1, rtx src2, rtx vccout, rtx vccin), GEN_VN (and,si3, A(rtx dest, rtx src1, rtx src2), A(dest, src1, src2)) GEN_VNM_NOEXEC (ds_bpermute,, A(rtx dest, rtx addr, rtx src, rtx exec), A(dest, addr, src, exec)) -GEN_VNM (gather,_expr, A(rtx dest, rtx addr, rtx as, rtx vol), - A(dest, addr, as, vol)) GEN_VN (sub,si3, A(rtx dest, rtx src1, rtx src2), A(dest, src1, src2)) GEN_VN_NOEXEC (vec_series,si, A(rtx dest, rtx x, rtx c), A(dest, x, c)) @@ -1471,11 +1469,11 @@ gcn_stepped_zero_int_parallel_p (rtx op, int step) /* {{{ Addresses, pointers and moves. */ /* Return true is REG is a valid place to store a pointer, - for instructions that require an SGPR. - FIXME rename. */ + for instructions that require an SGPR. Also check that the address + width is correct for the address space. */ static bool -gcn_address_register_p (rtx reg, machine_mode mode, bool strict) +gcn_scalar_address_register_p (rtx reg, machine_mode as_mode, bool strict) { if (GET_CODE (reg) == SUBREG) reg = SUBREG_REG (reg); @@ -1483,7 +1481,7 @@ gcn_address_register_p (rtx reg, machine_mode mode, bool strict) if (!REG_P (reg)) return false; - if (GET_MODE (reg) != mode) + if (GET_MODE (reg) != as_mode) return false; int regno = REGNO (reg); @@ -1504,10 +1502,11 @@ gcn_address_register_p (rtx reg, machine_mode mode, bool strict) } /* Return true is REG is a valid place to store a pointer, - for instructions that require a VGPR. */ + for instructions that require a VGPR. Also check that the address + width is correct for the address space. */ static bool -gcn_vec_address_register_p (rtx reg, machine_mode mode, bool strict) +gcn_vec_address_register_p (rtx reg, machine_mode as_mode, bool strict) { if (GET_CODE (reg) == SUBREG) reg = SUBREG_REG (reg); @@ -1515,7 +1514,12 @@ gcn_vec_address_register_p (rtx reg, machine_mode mode, bool strict) if (!REG_P (reg)) return false; - if (GET_MODE (reg) != mode) + /* Vector addresses are allowed, but MODE is always specified scalar + (otherwise the number of lanes would have to match too). */ + machine_mode scalar_reg_mode = (VECTOR_MODE_P (GET_MODE (reg)) + ? GET_MODE_INNER (GET_MODE (reg)) + : GET_MODE (reg)); + if (scalar_reg_mode != as_mode) return false; int regno = REGNO (reg); @@ -1534,25 +1538,64 @@ gcn_vec_address_register_p (rtx reg, machine_mode mode, bool strict) return VGPR_REGNO_P (regno); } +/* Return true if REG is a valid place to store a pointer, + for instructions that require an SGPR or VGPR, depending on the + mode of the address and the data type: + + Scalar load, scalar address: VGPR. + Vector load, scalar base: SGPR. (These get expanded later.) + Vector load, vector of addresses: VGPR. + + This function is only suitable for FLAT and GLOBAL address spaces. */ + +static bool +gcn_auto_address_register_p (rtx addr, machine_mode data_mode, bool strict) +{ + machine_mode addr_mode = GET_MODE (addr); + + if (!VECTOR_MODE_P (data_mode)) + return (!VECTOR_MODE_P (addr_mode) + && gcn_vec_address_register_p (addr, DImode, strict)); + else if (!VECTOR_MODE_P (addr_mode)) + return gcn_scalar_address_register_p (addr, DImode, strict); + else + return gcn_vec_address_register_p (addr, DImode, strict); +} + /* Return true if X would be valid inside a MEM using the Flat address space. */ bool -gcn_flat_address_p (rtx x, machine_mode mode) +gcn_flat_address_p (rtx x, machine_mode data_mode, bool strict) { - bool vec_mode = (GET_MODE_CLASS (mode) == MODE_VECTOR_INT - || GET_MODE_CLASS (mode) == MODE_VECTOR_FLOAT); - - if (vec_mode && gcn_address_register_p (x, DImode, false)) - return true; - - if (!vec_mode && gcn_vec_address_register_p (x, DImode, false)) + if (gcn_auto_address_register_p (x, data_mode, strict)) return true; if (GET_CODE (x) == PLUS - && gcn_vec_address_register_p (XEXP (x, 0), DImode, false) - && CONST_INT_P (XEXP (x, 1))) - return true; + && gcn_auto_address_register_p (XEXP (x, 0), data_mode, strict)) + { + rtx x1 = XEXP (x, 1); + + if (GET_CODE (x1) == VEC_DUPLICATE + && VECTOR_MODE_P (GET_MODE (x1)) + && GET_MODE_INNER (GET_MODE (x1)) == DImode) + x1 = XEXP (x1, 0); + else if (GET_CODE (x1) == CONST_VECTOR + && VECTOR_MODE_P (GET_MODE (x1)) + && GET_MODE_INNER (GET_MODE (x1)) == DImode + && gcn_constant_p (x1)) + x1 = XVECEXP (x1, 0, 0); + + if (GET_CODE (x1) == CONST_INT) + { + int offsetbits = (TARGET_11BIT_GLOBAL_OFFSET ? 11 : 12); + if (INTVAL (x1) >= 0 && INTVAL (x1) < (1 << offsetbits) + /* The low bits of the offset are ignored, even when + they're meant to realign the pointer. */ + && !(INTVAL (x1) & 0x3)) + return true; + } + } return false; } @@ -1561,15 +1604,34 @@ gcn_flat_address_p (rtx x, machine_mode mode) address space. */ bool -gcn_scalar_flat_address_p (rtx x) +gcn_scalar_flat_address_p (rtx x, bool strict) { - if (gcn_address_register_p (x, DImode, false)) + if (gcn_scalar_address_register_p (x, DImode, strict)) return true; if (GET_CODE (x) == PLUS - && gcn_address_register_p (XEXP (x, 0), DImode, false) - && CONST_INT_P (XEXP (x, 1))) - return true; + && gcn_scalar_address_register_p (XEXP (x, 0), DImode, strict)) + { + rtx x1 = XEXP (x, 1); + + /* FIXME: This is disabled because of the mode mismatch between + SImode (for the address or m0 register) and the DImode PLUS. + We'll need a zero_extend or similar. + + if (gcn_m0_register_p (x1, SImode, strict) + || gcn_scalar_address_register_p (x1, SImode, strict)) + return true; + else*/ + + if (GET_CODE (x1) == CONST_INT) + { + if (INTVAL (x1) >= 0 && INTVAL (x1) < (1 << 20) + /* The low bits of the offset are ignored, even when + they're meant to realign the pointer. */ + && !(INTVAL (x1) & 0x3)) + return true; + } + } return false; } @@ -1592,15 +1654,35 @@ gcn_scalar_flat_mem_p (rtx x) address spaces. */ bool -gcn_ds_address_p (rtx x) +gcn_ds_address_p (rtx x, bool strict = false) { - if (gcn_vec_address_register_p (x, SImode, false)) + if (gcn_vec_address_register_p (x, SImode, strict)) return true; if (GET_CODE (x) == PLUS - && gcn_vec_address_register_p (XEXP (x, 0), SImode, false) - && CONST_INT_P (XEXP (x, 1))) - return true; + && gcn_vec_address_register_p (XEXP (x, 0), SImode, strict)) + { + rtx x0 = XEXP (x, 0); + rtx x1 = XEXP (x, 1); + if (!gcn_vec_address_register_p (x0, DImode, strict)) + return false; + if (GET_CODE (x1) == REG) + { + if (GET_CODE (x1) != REG + || (REGNO (x1) <= FIRST_PSEUDO_REGISTER + && !gcn_ssrc_register_operand (x1, DImode))) + return false; + return true; + } + else if (GET_CODE (x1) == CONST_VECTOR + && GET_CODE (CONST_VECTOR_ELT (x1, 0)) == CONST_INT + && single_cst_vector_p (x1)) + { + x1 = CONST_VECTOR_ELT (x1, 0); + if (INTVAL (x1) >= 0 && INTVAL (x1) < (1 << 20)) + return true; + } + } return false; } @@ -1609,10 +1691,9 @@ gcn_ds_address_p (rtx x) address space. */ bool -gcn_global_address_p (rtx addr) +gcn_global_address_p (rtx addr, machine_mode data_mode, bool strict) { - if (gcn_address_register_p (addr, DImode, false) - || gcn_vec_address_register_p (addr, DImode, false)) + if (gcn_auto_address_register_p (addr, data_mode, strict)) return true; if (GET_CODE (addr) == PLUS) @@ -1624,23 +1705,46 @@ gcn_global_address_p (rtx addr) && INTVAL (offset) >= -(1 << offsetbits) && INTVAL (offset) < (1 << offsetbits)); - if ((gcn_address_register_p (base, DImode, false) - || gcn_vec_address_register_p (base, DImode, false)) + if (gcn_auto_address_register_p (base, data_mode, strict) && immediate_p) /* SGPR + CONST or VGPR + CONST */ return true; - if (gcn_address_register_p (base, DImode, false) + if (gcn_scalar_address_register_p (base, DImode, strict) && gcn_vgpr_register_operand (offset, SImode)) /* SPGR + VGPR */ return true; if (GET_CODE (base) == PLUS - && gcn_address_register_p (XEXP (base, 0), DImode, false) + && gcn_scalar_address_register_p (XEXP (base, 0), DImode, strict) && gcn_vgpr_register_operand (XEXP (base, 1), SImode) && immediate_p) /* (SGPR + VGPR) + CONST */ return true; + + bool vec_immediate_p = + (GET_CODE (offset) == CONST_VECTOR + && gcn_constant_p (offset) + /* Signed 12/13-bit immediate. */ + && INTVAL (CONST_VECTOR_ELT (offset, 0)) >= -(1 << offsetbits) + && INTVAL (CONST_VECTOR_ELT (offset, 0)) < (1 << offsetbits) + /* The low bits of the offset are ignored, even + when they're meant to realign the pointer. */ + && !(INTVAL (CONST_VECTOR_ELT (offset, 0)) & 0x3)); + + if (gcn_vec_address_register_p (base, DImode, strict) + && vec_immediate_p) + /* VGPR + CONST */ + return true; + + if (GET_CODE (base) == PLUS + && GET_CODE (XEXP (base, 0)) == VEC_DUPLICATE + && gcn_scalar_address_register_p (XEXP (XEXP (base, 0), 0), + DImode, strict) + && gcn_vgpr_register_operand (XEXP (base, 1), V64SImode) + && vec_immediate_p) + /* (SGPR + VGPR) + CONST */ + return true; } return false; @@ -1660,167 +1764,27 @@ static bool gcn_addr_space_legitimate_address_p (machine_mode mode, rtx x, bool strict, addr_space_t as, code_helper = ERROR_MARK) { + machine_mode addr_mode = GET_MODE (x); + if (VECTOR_MODE_P (addr_mode) + && (!VECTOR_MODE_P (mode) + || GET_MODE_NUNITS (mode) != GET_MODE_NUNITS (addr_mode))) + return false; + if (AS_SCALAR_FLAT_P (as)) { if (mode == QImode || mode == HImode) return 0; - - switch (GET_CODE (x)) - { - case REG: - return gcn_address_register_p (x, DImode, strict); - /* Addresses are in the form BASE+OFFSET - OFFSET is either 20bit unsigned immediate, SGPR or M0. - Writes and atomics do not accept SGPR. */ - case PLUS: - { - rtx x0 = XEXP (x, 0); - rtx x1 = XEXP (x, 1); - if (!gcn_address_register_p (x0, DImode, strict)) - return false; - /* FIXME: This is disabled because of the mode mismatch between - SImode (for the address or m0 register) and the DImode PLUS. - We'll need a zero_extend or similar. - - if (gcn_m0_register_p (x1, SImode, strict) - || gcn_address_register_p (x1, SImode, strict)) - return true; - else*/ - if (GET_CODE (x1) == CONST_INT) - { - if (INTVAL (x1) >= 0 && INTVAL (x1) < (1 << 20) - /* The low bits of the offset are ignored, even when - they're meant to realign the pointer. */ - && !(INTVAL (x1) & 0x3)) - return true; - } - return false; - } - - default: - break; - } + else + return gcn_scalar_flat_address_p (x, strict); } else if (AS_SCRATCH_P (as)) - return gcn_address_register_p (x, SImode, strict); + return gcn_scalar_address_register_p (x, SImode, strict); else if (AS_FLAT_P (as) || AS_FLAT_SCRATCH_P (as)) - { - if (GET_CODE (x) == REG) - return ((GET_MODE_CLASS (mode) == MODE_VECTOR_INT - || GET_MODE_CLASS (mode) == MODE_VECTOR_FLOAT) - ? gcn_address_register_p (x, DImode, strict) - : gcn_vec_address_register_p (x, DImode, strict)); - else - { - if (GET_CODE (x) == PLUS) - { - rtx x1 = XEXP (x, 1); - - if (VECTOR_MODE_P (mode) - ? !gcn_address_register_p (x, DImode, strict) - : !gcn_vec_address_register_p (x, DImode, strict)) - return false; - - if (GET_CODE (x1) == CONST_INT) - { - if (INTVAL (x1) >= 0 && INTVAL (x1) < (1 << 12) - /* The low bits of the offset are ignored, even when - they're meant to realign the pointer. */ - && !(INTVAL (x1) & 0x3)) - return true; - } - } - return false; - } - } + return gcn_flat_address_p (x, mode, strict); else if (AS_GLOBAL_P (as)) - { - if (GET_CODE (x) == REG) - return (gcn_address_register_p (x, DImode, strict) - || (!VECTOR_MODE_P (mode) - && gcn_vec_address_register_p (x, DImode, strict))); - else if (GET_CODE (x) == PLUS) - { - rtx base = XEXP (x, 0); - rtx offset = XEXP (x, 1); - - int offsetbits = (TARGET_11BIT_GLOBAL_OFFSET ? 11 : 12); - bool immediate_p = (GET_CODE (offset) == CONST_INT - /* Signed 12/13-bit immediate. */ - && INTVAL (offset) >= -(1 << offsetbits) - && INTVAL (offset) < (1 << offsetbits) - /* The low bits of the offset are ignored, even - when they're meant to realign the pointer. */ - && !(INTVAL (offset) & 0x3)); - - if (!VECTOR_MODE_P (mode)) - { - if ((gcn_address_register_p (base, DImode, strict) - || gcn_vec_address_register_p (base, DImode, strict)) - && immediate_p) - /* SGPR + CONST or VGPR + CONST */ - return true; - - if (gcn_address_register_p (base, DImode, strict) - && gcn_vgpr_register_operand (offset, SImode)) - /* SGPR + VGPR */ - return true; - - if (GET_CODE (base) == PLUS - && gcn_address_register_p (XEXP (base, 0), DImode, strict) - && gcn_vgpr_register_operand (XEXP (base, 1), SImode) - && immediate_p) - /* (SGPR + VGPR) + CONST */ - return true; - } - else - { - if (gcn_address_register_p (base, DImode, strict) - && immediate_p) - /* SGPR + CONST */ - return true; - } - } - else - return false; - } + return gcn_global_address_p (x, mode, strict); else if (AS_ANY_DS_P (as)) - switch (GET_CODE (x)) - { - case REG: - return (VECTOR_MODE_P (mode) - ? gcn_address_register_p (x, SImode, strict) - : gcn_vec_address_register_p (x, SImode, strict)); - /* Addresses are in the form BASE+OFFSET - OFFSET is either 20bit unsigned immediate, SGPR or M0. - Writes and atomics do not accept SGPR. */ - case PLUS: - { - rtx x0 = XEXP (x, 0); - rtx x1 = XEXP (x, 1); - if (!gcn_vec_address_register_p (x0, DImode, strict)) - return false; - if (GET_CODE (x1) == REG) - { - if (GET_CODE (x1) != REG - || (REGNO (x1) <= FIRST_PSEUDO_REGISTER - && !gcn_ssrc_register_operand (x1, DImode))) - return false; - } - else if (GET_CODE (x1) == CONST_VECTOR - && GET_CODE (CONST_VECTOR_ELT (x1, 0)) == CONST_INT - && single_cst_vector_p (x1)) - { - x1 = CONST_VECTOR_ELT (x1, 0); - if (INTVAL (x1) >= 0 && INTVAL (x1) < (1 << 20)) - return true; - } - return false; - } - - default: - break; - } + return gcn_ds_address_p (x, strict); else gcc_unreachable (); return false; @@ -1859,6 +1823,32 @@ gcn_addr_space_address_mode (addr_space_t addrspace) return gcn_addr_space_pointer_mode (addrspace); } +/* Implement TARGET_ADDR_SPACE_VALID_POINTER_MODE. + + Return true if MODE is an appropriate mode for ADDRSPACE. */ + +static bool +gcn_addr_space_valid_pointer_mode (machine_mode mode, addr_space_t addrspace) +{ + if (VECTOR_MODE_P (mode)) + mode = GET_MODE_INNER (mode); + + switch (addrspace) + { + case ADDR_SPACE_SCRATCH: + case ADDR_SPACE_LDS: + case ADDR_SPACE_GDS: + return mode == SImode; + case ADDR_SPACE_DEFAULT: + case ADDR_SPACE_FLAT: + case ADDR_SPACE_FLAT_SCRATCH: + case ADDR_SPACE_SCALAR_FLAT: + return mode == DImode; + default: + gcc_unreachable (); + } +} + /* Implement TARGET_ADDR_SPACE_SUBSET_P. Determine if one named address space is a subset of another. */ @@ -1995,37 +1985,85 @@ gcn_regno_mode_code_ok_for_base_p (int regno, return false; } -/* Implement MODE_CODE_BASE_REG_CLASS via gcn.h. +/* Shared code for MODE_CODE_BASE_REG_CLASS and INSN_BASE_REG_CLASS. */ - Return a suitable register class for memory addressing. */ - -reg_class -gcn_mode_code_base_reg_class (machine_mode mode, addr_space_t as, int oc, - int ic) +static reg_class +gcn_base_reg_class (machine_mode access_mode, addr_space_t as, + bool vector_base_p) { switch (as) { case ADDR_SPACE_DEFAULT: - return gcn_mode_code_base_reg_class (mode, DEFAULT_ADDR_SPACE, oc, ic); + return gcn_base_reg_class (access_mode, DEFAULT_ADDR_SPACE, vector_base_p); case ADDR_SPACE_SCALAR_FLAT: case ADDR_SPACE_SCRATCH: + gcc_assert (!vector_base_p); return SGPR_REGS; break; case ADDR_SPACE_FLAT: case ADDR_SPACE_FLAT_SCRATCH: case ADDR_SPACE_LDS: case ADDR_SPACE_GDS: - return ((GET_MODE_CLASS (mode) == MODE_VECTOR_INT - || GET_MODE_CLASS (mode) == MODE_VECTOR_FLOAT) - ? SGPR_REGS : VGPR_REGS); case ADDR_SPACE_GLOBAL: - return ((GET_MODE_CLASS (mode) == MODE_VECTOR_INT - || GET_MODE_CLASS (mode) == MODE_VECTOR_FLOAT) - ? SGPR_REGS : ALL_GPR_REGS); + if (vector_base_p) + return VGPR_REGS; + else if (GET_MODE_CLASS (access_mode) == MODE_VECTOR_INT + || GET_MODE_CLASS (access_mode) == MODE_VECTOR_FLOAT) + return SGPR_REGS; + else if (as == ADDR_SPACE_GLOBAL) + return ALL_GPR_REGS; + else + return VGPR_REGS; } gcc_unreachable (); } +/* Implement MODE_CODE_BASE_REG_CLASS via gcn.h. + + Return a suitable register class for memory addressing of scalar + address modes. This is only called when INSN_BASE_REG_CLASS can't. */ + +reg_class +gcn_mode_code_base_reg_class (machine_mode mode, addr_space_t as, int, int) +{ + return gcn_base_reg_class (mode, as, false); +} + +/* Implement INSN_BASE_REG_CLASS via gcn.h. + + Return a suitable register class for memory addressing when the INSN is + known. Supports both scalar and vector addressing modes. */ + +reg_class +gcn_insn_base_reg_class (rtx_insn *insn) +{ + gcc_assert (insn); + + const_rtx mem = NULL_RTX; + subrtx_iterator::array_type array; + FOR_EACH_SUBRTX (iter, array, PATTERN (insn), NONCONST) + { + if (MEM_P (*iter)) + { + mem = *iter; + + /* There might be multiple MEM in insn, but we don't know which + one the caller is focussed on. It's safer to select the first + vector one. */ + if (VECTOR_MODE_P (GET_MODE (XEXP (mem, 0)))) + break; + } + } + if (!mem) + return VGPR_REGS; + + machine_mode access_mode = GET_MODE (mem); + bool vector_base_p = VECTOR_MODE_P (GET_MODE (XEXP (mem, 0))); + addr_space_t as = MEM_ADDR_SPACE (mem); + + return gcn_base_reg_class (access_mode, as, vector_base_p); +} + /* Implement REGNO_OK_FOR_INDEX_P via gcn.h. Return true if REGNO is OK for index of memory addressing. */ @@ -2130,12 +2168,9 @@ gcn_expand_vector_init (rtx op0, rtx vec) if (mem_mask) { - emit_insn (gen_gathervNm_expr - (op0, gen_rtx_PLUS (addrmode, addr, - gen_rtx_VEC_DUPLICATE (addrmode, - const0_rtx)), - GEN_INT (DEFAULT_ADDR_SPACE), GEN_INT (0), - NULL, get_exec (mem_mask))); + rtx mem = gen_rtx_MEM (mode, addr); + emit_insn (gen_mov_exec (mode, op0, mem, gcn_gen_undef (mode), + get_exec (mem_mask))); prev = op0; initialized_mask = mem_mask; } @@ -2344,10 +2379,14 @@ gcn_expand_scalar_to_vector_address (machine_mode mode, rtx exec, rtx mem, gen_rtx_SIGN_EXTEND (pmode, tmplo)); } - return gen_rtx_PLUS (GET_MODE (new_base), new_base, - gen_rtx_VEC_DUPLICATE (GET_MODE (new_base), - (mem_index ? mem_index - : const0_rtx))); + rtx addr = gen_rtx_PLUS (GET_MODE (new_base), new_base, + gen_rtx_VEC_DUPLICATE (GET_MODE (new_base), + (mem_index ? mem_index + : const0_rtx))); + rtx newmem = gen_rtx_MEM (mode, addr); + set_mem_addr_space (newmem, MEM_ADDR_SPACE (mem)); + MEM_VOLATILE_P (newmem) = MEM_VOLATILE_P (mem); + return newmem; } /* Convert a BASE address, a vector of OFFSETS, and a SCALE, to addresses @@ -2403,6 +2442,21 @@ gcn_expand_scaled_offsets (addr_space_t as, rtx base, rtx offsets, rtx scale, gcc_unreachable (); } +/* Convert gather/scatter parameters to a vector MEM. */ + +rtx +gcn_gen_vector_mem (machine_mode mode, addr_space_t as, rtx scalarbase, + rtx vectoroffsets, rtx scale, bool unsigned_p, + bool volatile_p, rtx exec) +{ + rtx addr = gcn_expand_scaled_offsets (as, scalarbase, vectoroffsets, scale, + unsigned_p, exec); + rtx mem = gen_rtx_MEM (mode, addr); + set_mem_addr_space (mem, as); + MEM_VOLATILE_P (mem) = volatile_p; + return mem; +} + /* Return true if move from OP0 to OP1 is known to be executed in vector unit. */ @@ -2494,8 +2548,9 @@ gcn_secondary_reload (bool in_p, rtx x, reg_class_t rclass, case ADDR_SPACE_FLAT: case ADDR_SPACE_FLAT_SCRATCH: case ADDR_SPACE_GLOBAL: - if (GET_MODE_CLASS (reload_mode) == MODE_VECTOR_INT - || GET_MODE_CLASS (reload_mode) == MODE_VECTOR_FLOAT) + if ((GET_MODE_CLASS (reload_mode) == MODE_VECTOR_INT + || GET_MODE_CLASS (reload_mode) == MODE_VECTOR_FLOAT) + && !(MEM_P (x) && VECTOR_MODE_P (GET_MODE (XEXP (x, 0))))) { sri->icode = code_for_mov_sgprbase (reload_mode); break; @@ -2659,14 +2714,14 @@ gcn_valid_move_p (machine_mode mode, rtx dest, rtx src) if (MEM_P (dest) && AS_GLOBAL_P (MEM_ADDR_SPACE (dest)) - && (gcn_global_address_p (XEXP (dest, 0)) + && (gcn_global_address_p (XEXP (dest, 0), GET_MODE (dest)) || GET_CODE (XEXP (dest, 0)) == SYMBOL_REF || GET_CODE (XEXP (dest, 0)) == LABEL_REF) && gcn_vgpr_equivalent_register_operand (src, mode)) return true; else if (MEM_P (src) && AS_GLOBAL_P (MEM_ADDR_SPACE (src)) - && (gcn_global_address_p (XEXP (src, 0)) + && (gcn_global_address_p (XEXP (src, 0), GET_MODE (src)) || GET_CODE (XEXP (src, 0)) == SYMBOL_REF || GET_CODE (XEXP (src, 0)) == LABEL_REF) && gcn_vgpr_equivalent_register_operand (dest, mode)) @@ -3154,7 +3209,6 @@ move_callee_saved_registers (rtx sp, machine_function *offsets, rtx exec = gen_rtx_REG (DImode, EXEC_REG); rtx vcc = gen_rtx_REG (DImode, VCC_LO_REG); rtx offreg = gen_rtx_REG (SImode, SGPR_REGNO (22)); - rtx as = gen_rtx_CONST_INT (VOIDmode, STACK_ADDR_SPACE); HOST_WIDE_INT exec_set = 0; int offreg_set = 0; auto_vec saved_sgprs; @@ -3207,6 +3261,8 @@ move_callee_saved_registers (rtx sp, machine_function *offsets, gcn_operand_part (V64SImode, vsp, 1), const0_rtx, vcc, vcc, gcn_gen_undef (V64SImode), exec)); + rtx vspmem = gen_rtx_MEM (V64SImode, vsp); + set_mem_addr_space (vspmem, STACK_ADDR_SPACE); /* Move vectors. */ for (regno = FIRST_VGPR_REG, offset = 0; @@ -3231,9 +3287,9 @@ move_callee_saved_registers (rtx sp, machine_function *offsets, if (prologue) { - rtx insn = emit_insn (gen_scatterv64si_insn_1offset_exec - (vsp, const0_rtx, reg, as, const0_rtx, - exec)); + rtx insn = emit_insn (gen_movv64si_exec (vspmem, reg, + gcn_gen_undef (V64SImode), + exec)); /* Add CFI metadata. */ rtx note; @@ -3292,9 +3348,8 @@ move_callee_saved_registers (rtx sp, machine_function *offsets, add_reg_note (insn, REG_FRAME_RELATED_EXPR, note); } else - emit_insn (gen_gatherv64si_insn_1offset_exec - (reg, vsp, const0_rtx, as, const0_rtx, - gcn_gen_undef (V64SImode), exec)); + emit_insn (gen_movv64si_exec (reg, vspmem, gcn_gen_undef (V64SImode), + exec)); /* Move our VSP to the next stack entry. */ if (offreg_set != size) @@ -7230,6 +7285,7 @@ print_operand_address (FILE *file, rtx mem) rtx offset; addr_space_t as = MEM_ADDR_SPACE (mem); rtx addr = XEXP (mem, 0); + gcc_assert (REG_P (addr) || GET_CODE (addr) == PLUS); if (AS_SCRATCH_P (as)) @@ -7271,9 +7327,13 @@ print_operand_address (FILE *file, rtx mem) if (GET_CODE (base) == PLUS) { - /* (SGPR + VGPR) + CONST */ + /* (SGPR + VGPR) + CONST + Note: the offset is printed by %O. */ vgpr_offset = XEXP (base, 1); base = XEXP (base, 0); + + if (GET_CODE (base) == VEC_DUPLICATE) + base = XEXP (base, 0); } else { @@ -7309,8 +7369,6 @@ print_operand_address (FILE *file, rtx mem) output_operand_lossage ("bad ADDR_SPACE_GLOBAL address"); } } - else - output_operand_lossage ("bad ADDR_SPACE_GLOBAL address"); } else if (AS_ANY_DS_P (as)) switch (GET_CODE (addr)) @@ -7640,12 +7698,19 @@ print_operand (FILE *file, rtx x, int code) base = XEXP (x0, 0); if (GET_CODE (base) == PLUS) - /* (SGPR + VGPR) + CONST */ - /* Ignore the VGPR offset for this operand. */ - base = XEXP (base, 0); + { + /* (SGPR + VGPR) + CONST */ + /* Ignore the VGPR offset for this operand. */ + base = XEXP (base, 0); + + if (GET_CODE (base) == VEC_DUPLICATE) + base = XEXP (base, 0); + } + if (CONST_VECTOR_P (offset)) + offset = CONST_VECTOR_ELT (offset, 0); if (CONST_INT_P (offset)) - const_offset = XEXP (x0, 1); + const_offset = offset; else if (REG_P (offset)) /* SGPR + VGPR */ /* Ignore the VGPR offset for this operand. */ @@ -7684,6 +7749,8 @@ print_operand (FILE *file, rtx x, int code) rtx val = XEXP (x0, 1); if (GET_CODE (val) == CONST_VECTOR) val = CONST_VECTOR_ELT (val, 0); + if (GET_CODE (val) == VEC_DUPLICATE) + val = XEXP (val, 0); if (GET_CODE (val) != CONST_INT) { output_operand_lossage ("invalid %%xn code"); @@ -8097,6 +8164,8 @@ gcn_dwarf_register_span (rtx rtl) #define TARGET_ADDR_SPACE_ZERO_ADDRESS_VALID gcn_addr_space_zero_address_valid #undef TARGET_ADDR_SPACE_CONVERT #define TARGET_ADDR_SPACE_CONVERT gcn_addr_space_convert +#undef TARGET_ADDR_SPACE_VALID_POINTER_MODE +#define TARGET_ADDR_SPACE_VALID_POINTER_MODE gcn_addr_space_valid_pointer_mode #undef TARGET_ARG_PARTIAL_BYTES #define TARGET_ARG_PARTIAL_BYTES gcn_arg_partial_bytes #undef TARGET_ASM_ALIGNED_DI_OP diff --git a/gcc/config/gcn/gcn.h b/gcc/config/gcn/gcn.h index 87605edd79c..8582542e041 100644 --- a/gcc/config/gcn/gcn.h +++ b/gcc/config/gcn/gcn.h @@ -602,6 +602,7 @@ enum reg_class #define REGNO_REG_CLASS(REGNO) gcn_regno_reg_class (REGNO) #define MODE_CODE_BASE_REG_CLASS(MODE, AS, OUTER, INDEX) \ gcn_mode_code_base_reg_class (MODE, AS, OUTER, INDEX) +#define INSN_BASE_REG_CLASS(INSN) gcn_insn_base_reg_class (INSN) #define REGNO_MODE_CODE_OK_FOR_BASE_P(NUM, MODE, AS, OUTER, INDEX) \ gcn_regno_mode_code_ok_for_base_p (NUM, MODE, AS, OUTER, INDEX) #define INDEX_REG_CLASS VGPR_REGS diff --git a/gcc/config/gcn/gcn.md b/gcc/config/gcn/gcn.md index f95e7e555c9..4c56b1b5e96 100644 --- a/gcc/config/gcn/gcn.md +++ b/gcc/config/gcn/gcn.md @@ -67,6 +67,7 @@ (define_c_enum "unspec" [ UNSPEC_VECTOR UNSPEC_BPERMUTE UNSPEC_SGPRBASE + UNSPEC_MASKLOAD UNSPEC_MEMORY_BARRIER UNSPEC_SMIN_DPP_SHR UNSPEC_SMAX_DPP_SHR UNSPEC_UMIN_DPP_SHR UNSPEC_UMAX_DPP_SHR @@ -81,8 +82,6 @@ (define_c_enum "unspec" [ UNSPEC_CMUL_ADD UNSPEC_CMUL_SUB UNSPEC_CADD90 UNSPEC_CADD270 - UNSPEC_GATHER - UNSPEC_SCATTER UNSPEC_RCP UNSPEC_FLBIT_INT UNSPEC_FLOOR UNSPEC_CEIL UNSPEC_SIN UNSPEC_COS UNSPEC_EXP2 UNSPEC_LOG2 @@ -1976,7 +1975,7 @@ (define_insn "atomic_fetch_" "0 /* Disabled. */" "@ s_atomic_\t%0, %1, %2 glc\;s_waitcnt\tlgkmcnt(0) - flat_atomic_\t%0, %1, %2 %G2\;s_waitcnt\t0 + flat_atomic_\t%0, %1, %2%O1 %G2\;s_waitcnt\t0 global_atomic_\t%0, %A1, %2%O1 %G2\;s_waitcnt\tvmcnt(0)" [(set_attr "type" "smem,flat,flat") (set_attr "flatmemaccess" "*,atomicwait,atomicwait") @@ -1998,7 +1997,7 @@ (define_insn "atomic_" "0 /* Disabled. */" "@ s_atomic_\t%0, %1\;s_waitcnt\tlgkmcnt(0) - flat_atomic_\t%0, %1\;s_waitcnt\t0 + flat_atomic_\t%0, %1%O0\;s_waitcnt\t0 global_atomic_\t%A0, %1%O0\;s_waitcnt\tvmcnt(0)" [(set_attr "type" "smem,flat,flat") (set_attr "flatmemaccess" "*,atomicwait,atomicwait") @@ -2045,7 +2044,7 @@ (define_insn "sync_compare_and_swap_insn" "" "@ s_atomic_cmpswap\t%0, %1, %2 glc\;s_waitcnt\tlgkmcnt(0) - flat_atomic_cmpswap\t%0, %1, %2 %G2\;s_waitcnt\t0 + flat_atomic_cmpswap\t%0, %1, %2%O1 %G2\;s_waitcnt\t0 global_atomic_cmpswap\t%0, %A1, %2%O1 %G2\;s_waitcnt\tvmcnt(0)" [(set_attr "type" "smem,flat,flat") (set_attr "length" "12") From patchwork Wed Jul 8 14:46:33 2026 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Andrew Stubbs X-Patchwork-Id: 138757 Return-Path: X-Original-To: patchwork@sourceware.org Delivered-To: patchwork@sourceware.org Received: from vm01.sourceware.org (localhost [IPv6:::1]) by sourceware.org (Postfix) with ESMTP id F21834BA2E18 for ; Wed, 8 Jul 2026 14:49:21 +0000 (GMT) DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org F21834BA2E18 Authentication-Results: sourceware.org; dkim=pass (2048-bit key, secure) header.d=baylibre.com header.i=@baylibre.com header.a=rsa-sha256 header.s=google header.b=Y5aoOKb/ X-Original-To: gcc-patches@gcc.gnu.org Delivered-To: gcc-patches@gcc.gnu.org Received: from mail-wm1-x32e.google.com (mail-wm1-x32e.google.com [IPv6:2a00:1450:4864:20::32e]) by sourceware.org (Postfix) with ESMTPS id E67974BA2E05 for ; Wed, 8 Jul 2026 14:46:48 +0000 (GMT) DMARC-Filter: OpenDMARC Filter v1.4.2 sourceware.org E67974BA2E05 Authentication-Results: sourceware.org; dmarc=none (p=none dis=none) header.from=baylibre.com Authentication-Results: sourceware.org; spf=pass smtp.mailfrom=baylibre.com ARC-Filter: OpenARC Filter v1.0.0 sourceware.org E67974BA2E05 Authentication-Results: sourceware.org; arc=none smtp.remote-ip=2a00:1450:4864:20::32e ARC-Seal: i=1; a=rsa-sha256; d=sourceware.org; s=key; t=1783522009; cv=none; b=wCFZDlDW/c+vZomM6eiLpS1ygxeOjtTbJF2NfT1gKs+HJRph5KjGYaNjM19879MCNy2LOov2wj87q9+XWx/9r610S8zFd/zEJrRZvbDskZnZCLcA/3JTtki2RfU8xYJzPpy+m3tbabOkdU7TZ0i9vv+FwB85TmpMIjG5u87lPv8= ARC-Message-Signature: i=1; a=rsa-sha256; d=sourceware.org; s=key; t=1783522009; c=relaxed/simple; bh=NTFVxpgA68LrJKyuSv+SgPC+eCfYOipuhOK0zfNHFPQ=; h=DKIM-Signature:From:To:Subject:Date:Message-ID:MIME-Version; b=IMfvFGbV5if6eZKpk8PImvf7p4i2H1wG++H/WzAPezt8wSW/zVI+Fzw7HgTilH8/Ms+ZOULZQlYljcuS7F3zFl5XV4FkdHC9RhXVy4WVJl0Rhy9Sx4h9ITt5Sdxrb5rJgTJKFdA/E7c669ViaGB4WbKdxTuQqSh9YNG3I4cpL1M= ARC-Authentication-Results: i=1; sourceware.org; dkim=pass (2048-bit key, secure) header.d=baylibre.com header.i=@baylibre.com header.a=rsa-sha256 header.s=google header.b=Y5aoOKb/ DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org E67974BA2E05 Received: by mail-wm1-x32e.google.com with SMTP id 5b1f17b1804b1-490cf3000f0so5476545e9.1 for ; Wed, 08 Jul 2026 07:46:48 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=baylibre.com; s=google; t=1783522008; x=1784126808; darn=gcc.gnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:to:from:from:to:cc:subject:date:message-id :reply-to:content-type; bh=moS302CW4i84Z0NnuPkuA1twVdjtVOLIGaoih1M1BrE=; b=Y5aoOKb/kpo6xLFLmU1Zxvu37NLsEoLZHI680h91YMRDUc2m6VGqCFZokpl+3zumFL Fg/Y5ymwSZ3ibxXiOldK/cxmT/gPrGHmIseLpSSqIzqdGzv6DdzJ0IRRb4/69mfQLUvB Tg1p/RK1Sk/BsmQF8AUSBhul6xJYCuuGw4pMIO4AOizkgw5naYCGWWsMLC+hIOrcpx3J E60aLblTMGzkWES0HOp35fl4B+kGvSex5ETY7PNLj/6zplnjfenAN1Jc9FrfLw4n+d0N M7erx6rR6RgTYWZCwdt0G6tu9svhNISMVxuI7F4Ft5jXtPkGsjBm+VmZXN/r1D+UHvU5 MHxQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783522008; x=1784126808; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:to:from:x-gm-gg:x-gm-message-state:from:to :cc:subject:date:message-id:reply-to:content-type; bh=moS302CW4i84Z0NnuPkuA1twVdjtVOLIGaoih1M1BrE=; b=VlXLZ7MBGjyVlHkyzBP/W3TwvS81y8phakdESUpjaa3epNqAFQYHsogLSYSRCb5+O8 coSHyNV7Br8M/KsM9lcJcFQ3xEccqHpUuX7bkBdSWxlHmYvyB7xAb69l3ZK/+jsdsJl9 S45RpkKSe/9eoe9KGG4eZfO2vswqCvgJKWbEcnTHOJxy+qowCF15BbB3F1yRMKg64O+t vwooVaqZoCPAAjE3azg1o8a5tYHmJMxx0J2dlRT8RC7SOog4wmo2w6A7W4NDNM2htgQk AalCDQVNqNef7papJbeB5CQx/ZwqYh1U2V6n3SgaAoUbFA7WiZrzUXwy2t3pBm8L9I7d 0CWg== X-Gm-Message-State: AOJu0YzMKsdJEmptNSRwJLsHni4l6mT2vEK0zlBsj80a2e7VwoVuWNjr BGNQgv3prAhA+62p8iDwu64h6bfiAfViAMDjR4yXQBnvpWFWy1C33ifJCnDUBL73WhtWzC/33Rz eDJer X-Gm-Gg: AfdE7cm5Vt8+xWQWN9KDdl3u6CKOKGYvKqFIexBC1MC3Wpt6LTg0Lo5rb9BCeaUceDz Tcy7BGZPhHeXKuWHSVPbCLHueKX98aNK1bdknfipHckgNCuBmpcwv8+8SY2BpCBHOkfk0CaEa+K Nw9klY3V84fwgK27kjMzfv3gzQF4e0sInVWLwlUkITWiqnXCSY/a33TOebeRMI1Xe6uD30F8rab 0jXvBdeuRLoikLQra7DE5wu3gjqdc1zZiVdK22aBL6htMgRN2wgg4ecVxpkmkJgEOTcekrS86Cw hrxeVXR4gPKSiZ7jtLHllACpf5aHNriiYEH5R1SFy+/KwmakQT2k9qK3bDpvHCeDN+M8pda4GJK j4LCl+acMl74KXuaH/6qwXqxzTXecFGw4/4+jaQr97mPKPtOaCKBaQl3RV9rgR8Lmmgekz3A3ab FrsBSvMhc1CJPZm7xZVOoMST6gIw== X-Received: by 2002:a05:600d:84ca:20b0:493:c984:db9c with SMTP id 5b1f17b1804b1-493e6898b54mr22011365e9.2.1783522007526; Wed, 08 Jul 2026 07:46:47 -0700 (PDT) Received: from vbuild-02.baylibre ([217.13.61.132]) by smtp.googlemail.com with ESMTPSA id ffacd0b85a97d-47a9e4d83bdsm43362109f8f.13.2026.07.08.07.46.46 for (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 08 Jul 2026 07:46:47 -0700 (PDT) From: Andrew Stubbs To: gcc-patches@gcc.gnu.org Subject: [PATCH 3/3] amdgcn: Add vector atomics Date: Wed, 8 Jul 2026 14:46:33 +0000 Message-ID: <20260708144633.1530935-4-ams@baylibre.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260708144633.1530935-1-ams@baylibre.com> References: <20260708144633.1530935-1-ams@baylibre.com> MIME-Version: 1.0 X-Spam-Status: No, score=-11.3 required=5.0 tests=BAYES_00, DKIM_SIGNED, DKIM_VALID, DKIM_VALID_AU, DKIM_VALID_EF, GIT_PATCH_0, RCVD_IN_DNSWL_NONE, SPF_HELO_NONE, SPF_PASS, TXREP shortcircuit=no autolearn=ham autolearn_force=no version=3.4.6 X-Spam-Checker-Version: SpamAssassin 3.4.6 (2021-04-09) on sourceware.org X-BeenThere: gcc-patches@gcc.gnu.org X-Mailman-Version: 2.1.30 Precedence: list List-Id: Gcc-patches mailing list List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: gcc-patches-bounces~patchwork=sourceware.org@gcc.gnu.org This patch utilizes the new "(mem (reg:))" support to add vector atomics. These work exactly like the scalar atomic operation, but 64 times in parallel. gcc/ChangeLog: * config/gcn/gcn.md (UNSPEC_PAIR): New. (X): Add vector modes. (ATOMICMODE): New mode iterator. (atomic_fetch_): Use ATOMICMODE iterator. (atomic_): Likewise. (x2): Add vector modes. (size): Delete. (bitsize): Add vector modes. (createpair): New insn. (sync_compare_and_swap): Use ATOMICMODE iterator. (sync_compare_and_swap_insn): Likewise. (sync_compare_and_swap_lds_insn): Likewise. (atomic_load): Likewise. (atomic_store): Likewise. (atomic_exchange): Likewise. --- gcc/config/gcn/gcn.md | 114 +++++++++++++++++++++++++----------------- 1 file changed, 69 insertions(+), 45 deletions(-) diff --git a/gcc/config/gcn/gcn.md b/gcc/config/gcn/gcn.md index 4c56b1b5e96..75949f64dea 100644 --- a/gcc/config/gcn/gcn.md +++ b/gcc/config/gcn/gcn.md @@ -86,7 +86,8 @@ (define_c_enum "unspec" [ UNSPEC_FLBIT_INT UNSPEC_FLOOR UNSPEC_CEIL UNSPEC_SIN UNSPEC_COS UNSPEC_EXP2 UNSPEC_LOG2 UNSPEC_LDEXP UNSPEC_FREXP_EXP UNSPEC_FREXP_MANT - UNSPEC_DIV_SCALE UNSPEC_DIV_FMAS UNSPEC_DIV_FIXUP]) + UNSPEC_DIV_SCALE UNSPEC_DIV_FMAS UNSPEC_DIV_FIXUP + UNSPEC_PAIR]) ;; }}} ;; {{{ Attributes @@ -1921,7 +1922,11 @@ (define_expand "ti3" ; the programmer to get this right. (define_code_iterator atomicops [plus minus and ior xor]) -(define_mode_attr X [(SI "") (DI "_X2")]) +(define_mode_attr X [(SI "") (V64SI "") + (DI "_X2") (V64DI "_X2")]) + +(define_mode_iterator ATOMICMODE + [SI DI V64SI V64DI]) ;; TODO compare_and_swap test_and_set inc dec ;; Hardware also supports min and max, but GCC does not. @@ -1963,13 +1968,13 @@ (define_insn "*memory_barrier" ; reliably - they can cause hangs or incorrect results. ; TODO: flush caches according to memory model (define_insn "atomic_fetch_" - [(set (match_operand:SIDI 0 "register_operand" "=Sm, v, v") - (match_operand:SIDI 1 "memory_operand" "+RS,RF,RM")) + [(set (match_operand:ATOMICMODE 0 "register_operand" "=Sm, v, v") + (match_operand:ATOMICMODE 1 "memory_operand" "+RS,RfRF,RmRM")) (set (match_dup 1) - (unspec_volatile:SIDI - [(atomicops:SIDI + (unspec_volatile:ATOMICMODE + [(atomicops:ATOMICMODE (match_dup 1) - (match_operand:SIDI 2 "register_operand" " Sm, v, v"))] + (match_operand:ATOMICMODE 2 "register_operand" " Sm, v, v"))] UNSPECV_ATOMIC)) (use (match_operand 3 "const_int_operand"))] "0 /* Disabled. */" @@ -1987,11 +1992,11 @@ (define_insn "atomic_fetch_" ; you might expect from a concurrent non-atomic read-modify-write. ; TODO: flush caches according to memory model (define_insn "atomic_" - [(set (match_operand:SIDI 0 "memory_operand" "+RS,RF,RM") - (unspec_volatile:SIDI - [(atomicops:SIDI + [(set (match_operand:ATOMICMODE 0 "memory_operand" "+RS,RfRF,RmRM") + (unspec_volatile:ATOMICMODE + [(atomicops:ATOMICMODE (match_dup 0) - (match_operand:SIDI 1 "register_operand" " Sm, v, v"))] + (match_operand:ATOMICMODE 1 "register_operand" " Sm, v, v"))] UNSPECV_ATOMIC)) (use (match_operand 2 "const_int_operand"))] "0 /* Disabled. */" @@ -2003,15 +2008,35 @@ (define_insn "atomic_" (set_attr "flatmemaccess" "*,atomicwait,atomicwait") (set_attr "length" "12")]) -(define_mode_attr x2 [(SI "DI") (DI "TI")]) -(define_mode_attr size [(SI "4") (DI "8")]) -(define_mode_attr bitsize [(SI "32") (DI "64")]) +(define_mode_attr x2 [(SI "DI") (DI "TI") (V64SI "V64DI") (V64DI "V64TI")]) +(define_mode_attr bitsize [(SI "32") (DI "64") (V64SI "32") (V64DI "64")]) + +(define_insn_and_split "createpair" + [(set (match_operand: 0 "register_operand" "=&v,&Sm") + (unspec: [(match_operand:ATOMICMODE 1 "register_operand" "v, Sm") + (match_operand:ATOMICMODE 2 "register_operand" "v, Sm")] + UNSPEC_PAIR))] + "" + "#" + "reload_completed" + [(const_int 0)] + { + int parts = / 32; + int outpart = 0; + for (int inpart = 0; inpart < parts; inpart++, outpart++) + emit_move_insn (gcn_operand_part (mode, operands[0], outpart), + gcn_operand_part (mode, operands[1], inpart)); + for (int inpart = 0; inpart < parts; inpart++, outpart++) + emit_move_insn (gcn_operand_part (mode, operands[0], outpart), + gcn_operand_part (mode, operands[2], inpart)); + DONE; + }) (define_expand "sync_compare_and_swap" - [(match_operand:SIDI 0 "register_operand") - (match_operand:SIDI 1 "memory_operand") - (match_operand:SIDI 2 "register_operand") - (match_operand:SIDI 3 "register_operand")] + [(match_operand:ATOMICMODE 0 "register_operand") + (match_operand:ATOMICMODE 1 "memory_operand") + (match_operand:ATOMICMODE 2 "register_operand") + (match_operand:ATOMICMODE 3 "register_operand")] "" { if (MEM_ADDR_SPACE (operands[1]) == ADDR_SPACE_LDS) @@ -2024,22 +2049,21 @@ (define_expand "sync_compare_and_swap" } /* Operands 2 and 3 must be placed in consecutive registers, and passed - as a combined value. */ + as a combined value. Subregs would work for the scalar case, but + not for the vector case. */ rtx src_cmp = gen_reg_rtx (mode); - emit_move_insn (gen_rtx_SUBREG (mode, src_cmp, 0), operands[3]); - emit_move_insn (gen_rtx_SUBREG (mode, src_cmp, ), operands[2]); - emit_insn (gen_sync_compare_and_swap_insn (operands[0], - operands[1], + emit_insn (gen_createpair (src_cmp, operands[3], operands[2])); + emit_insn (gen_sync_compare_and_swap_insn (operands[0], operands[1], src_cmp)); DONE; }) (define_insn "sync_compare_and_swap_insn" - [(set (match_operand:SIDI 0 "register_operand" "=Sm, v, v") - (match_operand:SIDI 1 "memory_operand" "+RS,RF,RM")) + [(set (match_operand:ATOMICMODE 0 "register_operand" "=Sm, v, v") + (match_operand:ATOMICMODE 1 "memory_operand" "+RS,RfRF,RmRM")) (set (match_dup 1) - (unspec_volatile:SIDI - [(match_operand: 2 "register_operand" " Sm, v, v")] + (unspec_volatile:ATOMICMODE + [(match_operand: 2 "register_operand" " Sm, v, v")] UNSPECV_ATOMIC))] "" "@ @@ -2051,14 +2075,14 @@ (define_insn "sync_compare_and_swap_insn" (set_attr "flatmemaccess" "*,cmpswapx2,cmpswapx2")]) (define_insn "sync_compare_and_swap_lds_insn" - [(set (match_operand:SIDI 0 "register_operand" "= v") - (unspec_volatile:SIDI - [(match_operand:SIDI 1 "memory_operand" "+RL")] + [(set (match_operand:ATOMICMODE 0 "register_operand" "= v") + (unspec_volatile:ATOMICMODE + [(match_operand:ATOMICMODE 1 "memory_operand" "+RLRl")] UNSPECV_ATOMIC)) (set (match_dup 1) - (unspec_volatile:SIDI - [(match_operand:SIDI 2 "register_operand" " v") - (match_operand:SIDI 3 "register_operand" " v")] + (unspec_volatile:ATOMICMODE + [(match_operand:ATOMICMODE 2 "register_operand" " v") + (match_operand:ATOMICMODE 3 "register_operand" " v")] UNSPECV_ATOMIC))] "" { @@ -2071,11 +2095,11 @@ (define_insn "sync_compare_and_swap_lds_insn" (set_attr "length" "12")]) (define_insn "atomic_load" - [(set (match_operand:SIDI 0 "register_operand" "=Sm, v, v") - (unspec_volatile:SIDI - [(match_operand:SIDI 1 "memory_operand" " RS,RF,RM")] + [(set (match_operand:ATOMICMODE 0 "register_operand" "=Sm, v, v") + (unspec_volatile:ATOMICMODE + [(match_operand:ATOMICMODE 1 "memory_operand" " RS,RfRF,RmRM")] UNSPECV_ATOMIC)) - (use (match_operand:SIDI 2 "immediate_operand" " i, i, i"))] + (use (match_operand:SI 2 "immediate_operand" " i, i, i"))] "" { /* FIXME: RDNA cache instructions may be too conservative? */ @@ -2173,11 +2197,11 @@ (define_insn "atomic_load" (set_attr "rdna" "no,*,*")]) (define_insn "atomic_store" - [(set (match_operand:SIDI 0 "memory_operand" "=RS,RF,RM") - (unspec_volatile:SIDI - [(match_operand:SIDI 1 "register_operand" " Sm, v, v")] + [(set (match_operand:ATOMICMODE 0 "memory_operand" "=RS,RfRF,RmRM") + (unspec_volatile:ATOMICMODE + [(match_operand:ATOMICMODE 1 "register_operand" " Sm, v, v")] UNSPECV_ATOMIC)) - (use (match_operand:SIDI 2 "immediate_operand" " i, i, i"))] + (use (match_operand:SI 2 "immediate_operand" " i, i, i"))] "" { switch (INTVAL (operands[2])) @@ -2260,11 +2284,11 @@ (define_insn "atomic_store" (set_attr "rdna" "no,*,*")]) (define_insn "atomic_exchange" - [(set (match_operand:SIDI 0 "register_operand" "=Sm, v, v") - (match_operand:SIDI 1 "memory_operand" "+RS,RF,RM")) + [(set (match_operand:ATOMICMODE 0 "register_operand" "=Sm, v, v") + (match_operand:ATOMICMODE 1 "memory_operand" "+RS,RfRF,RmRM")) (set (match_dup 1) - (unspec_volatile:SIDI - [(match_operand:SIDI 2 "register_operand" " Sm, v, v")] + (unspec_volatile:ATOMICMODE + [(match_operand:ATOMICMODE 2 "register_operand" " Sm, v, v")] UNSPECV_ATOMIC)) (use (match_operand 3 "immediate_operand"))] ""