From patchwork Wed Jul 8 06:04:36 2026 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Kyrylo Tkachov X-Patchwork-Id: 138709 Return-Path: X-Original-To: patchwork@sourceware.org Delivered-To: patchwork@sourceware.org Received: from vm01.sourceware.org (localhost [IPv6:::1]) by sourceware.org (Postfix) with ESMTP id AB4B14BA2E11 for ; Wed, 8 Jul 2026 06:06:28 +0000 (GMT) DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org AB4B14BA2E11 Authentication-Results: sourceware.org; dkim=pass (2048-bit key, unprotected) header.d=Nvidia.com header.i=@Nvidia.com header.a=rsa-sha256 header.s=selector2 header.b=aYvc3KPn X-Original-To: gcc-patches@gcc.gnu.org Delivered-To: gcc-patches@gcc.gnu.org Received: from DM1PR04CU001.outbound.protection.outlook.com (mail-centralusazon11010039.outbound.protection.outlook.com [52.101.61.39]) by sourceware.org (Postfix) with ESMTPS id 71EAD4BA2E24 for ; Wed, 8 Jul 2026 06:05:20 +0000 (GMT) DMARC-Filter: OpenDMARC Filter v1.4.2 sourceware.org 71EAD4BA2E24 Authentication-Results: sourceware.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com Authentication-Results: sourceware.org; spf=fail smtp.mailfrom=nvidia.com ARC-Filter: OpenARC Filter v1.0.0 sourceware.org 71EAD4BA2E24 Authentication-Results: sourceware.org; arc=pass smtp.remote-ip=52.101.61.39 ARC-Seal: i=2; a=rsa-sha256; d=sourceware.org; s=key; t=1783490726; cv=pass; b=W3fsCzCpkqshZZgJT/Z7GtphQ34meexWNGP6n9LufyXac0GNnmwMqnTw2qIH9dJZy+kgAGDxw5eWXj1Dc7d8GhszfCSz/aGy4yu4r5cLiLKQUn1EbxAsbtaZB1Qwlea3WRfoAxv1Mj9+adOXvpgRIGvHpblE7fIP8zAzohBQGrU= ARC-Message-Signature: i=2; a=rsa-sha256; d=sourceware.org; s=key; t=1783490726; c=relaxed/simple; bh=IGeOkoINzj/xnHXYW3/UKMX+ZHii93XOraA1BEJ4JXg=; h=DKIM-Signature:From:To:Subject:Date:Message-ID:MIME-Version; b=tqHuVld+Nz4oVPOMaXSSEKAYLXo1bif3zIWfaAKxTlyyWh2bZKyPXmV6Ee+oDznUPu+E35/mMjlsrqMnLBfkE7SCaiWOpWI4VZbkYe/TrGb9CG22YcbvA9wTZpmR1vJPy/71m90Y6Y6q6+nJMV8RKoaaHOTRJCQ6JTP9YE85QQ8= ARC-Authentication-Results: i=2; sourceware.org; dkim=pass (2048-bit key, unprotected) header.d=Nvidia.com header.i=@Nvidia.com header.a=rsa-sha256 header.s=selector2 header.b=aYvc3KPn DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org 71EAD4BA2E24 ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=LtBFWZSCpStWxcFMgvQdW1Bv2T/ePvTM0elrU9p1yVKuB78wIIpehPABkkcVJn4a9Yhw9LYDbdjBkbhoAHQGmPn3K3rbSURvrLh6WqbeJfX24Rid/44y7jTz/7T244cZwqAfvbsqmLFQaUSq5B4H7faenw6fCf44C7StizeSfLRMdaSy8P5u8fpZ8EkzC03ssRyv34nSjIVl5DGMkt/kouxx3BqtYOULtrXhopnNqw2nbLeswH1ZIMkHL5qmQiRPnSxemIQizQJ20vTKheUw+gOrGylG2xQPYvvD5PhhH28Q1DQJy0PXe9GoR5iGsG55Ylq7i55T3Cvyu+5UafC2cA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=XnOxwreYxSVKUAA4lUqL3NFpedZNrgH4t+92OThVK9Q=; b=Y+a4OiaSPGg8vWAIknWCgXeLHFClnvVcre4ApLPqMs2p8F74xyFRrMgWwtVY19nUr8nJa/P3+BbFzRsL4BYrDuO5pOi0seVEj/9OOuVF2QuP+CS+Tk+ZdeoUSdiw3hDWUYrgo4wDmHJy2GG8RPuKKCQKU1K0PlWQE58lyZFBBi55Xbe6cXColqa+NmFjS1aVFbd5NfwmrYMmcopCnhPNmiPB7mGvkdyD0Ddo2xnehiPFZ8HsWsRD0ev8RvkQIb+uqbkwZALk1a4BWhGja2rBzYLHa/MsL9siWv7j2Ptw16AweD4lZYBjw+NK6jgsZTM1JeejzfUjcs7EQBZ9XqelzA== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 216.228.117.160) smtp.rcpttodomain=gcc.gnu.org smtp.mailfrom=nvidia.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=nvidia.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=Nvidia.com; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=XnOxwreYxSVKUAA4lUqL3NFpedZNrgH4t+92OThVK9Q=; b=aYvc3KPnZ9unqpQsw8z00DC1CIK3ofOP8PW4Bu8QndFSElNYqNCd5pn/7YqXEdRVrlbnjcp8FrFuZpQPlDZlzAVbpSCSGCD1HZQamm2OcWVL4FkNUfLBEbhJDu070TNVSRiJPHqgZAOFbrIzqZJZCMU0qxrkuD0m+Slu/LeRgIUe7z8bZuGMQLnmYSU4ydrQa41+GyRQWouF+WG9GHlqzdLR3k0xxTSsZn79WQnEtpcKxD3S12/7GdMxdRKbkaWYZOls1S6dUEXx2sjwiBcnF7+s4QXF/3X91clawxfMnScn2a1q3oEJmI4/ADBYSaMcFZi9C3ic/HyST9LSxSIdJA== Received: from DS1P221CA0007.NAMP221.PROD.OUTLOOK.COM (2603:10b6:8:451::18) by SJ1PR12MB6100.namprd12.prod.outlook.com (2603:10b6:a03:45d::20) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.181.10; Wed, 8 Jul 2026 06:05:13 +0000 Received: from DS1PEPF00017095.namprd03.prod.outlook.com (2603:10b6:8:451:cafe::2c) by DS1P221CA0007.outlook.office365.com (2603:10b6:8:451::18) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.202.10 via Frontend Transport; Wed, 8 Jul 2026 06:05:12 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 216.228.117.160) smtp.mailfrom=nvidia.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=nvidia.com; Received-SPF: Pass (protection.outlook.com: domain of nvidia.com designates 216.228.117.160 as permitted sender) receiver=protection.outlook.com; client-ip=216.228.117.160; helo=mail.nvidia.com; pr=C Received: from mail.nvidia.com (216.228.117.160) by DS1PEPF00017095.mail.protection.outlook.com (10.167.17.138) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.181.6 via Frontend Transport; Wed, 8 Jul 2026 06:05:12 +0000 Received: from rnnvmail201.nvidia.com (10.129.68.8) by mail.nvidia.com (10.129.200.66) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.20; Tue, 7 Jul 2026 23:04:54 -0700 Received: from ktkachov-mlt.nvidia.com (10.126.231.37) by rnnvmail201.nvidia.com (10.129.68.8) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.20; Tue, 7 Jul 2026 23:04:52 -0700 From: To: CC: , , Kyrylo Tkachov Subject: [PATCH 1/2] tree-optimization/126043 - form comparison masks in early phiopt Date: Wed, 8 Jul 2026 08:04:36 +0200 Message-ID: <20260708060437.26491-1-ktkachov@nvidia.com> X-Mailer: git-send-email 2.50.1 MIME-Version: 1.0 X-Originating-IP: [10.126.231.37] X-ClientProxiedBy: rnnvmail202.nvidia.com (10.129.68.7) To rnnvmail201.nvidia.com (10.129.68.8) X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: DS1PEPF00017095:EE_|SJ1PR12MB6100:EE_ X-MS-Office365-Filtering-Correlation-Id: b8559cb8-10f7-4a87-bf03-08dedcb6dfa9 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0; ARA:13230040|376014|82310400026|36860700016|1800799024|23010399003|5023799004|11063799006|56012099006|3023799007|18002099003|6133799003|13003099007; X-Microsoft-Antispam-Message-Info: /eFYpBlJyxVLVF4INtR7ZT98Phajs//+sYV6ICNl8mHz1ZS0o09xyBmcXI+Dpr/XJdCpFN5PTuk5p6yGMZBjZgC4mcKZMlfvJtzFNTCf0fCws61NCcSNN1IQZDmgX8Kp6XZejj4mBGkWoddgpXeWxrc094OiaW22jmOyngqfzA+g8ZWre3doN+5ouBWpvUiFFWWPbcMdqyanwLwWbPMeGAB/yUOp91e7suDV/h7wEAcDOlYvRS9Jzpg6aw13Xt4wDxFG23JhAVcru/oJtPIfDMrRy0tcyd79CCi+vqgz5KKCAScmp9h5XOc8IOunWcYlSsTWeBAMB3Ra6rnOIvY5ZaHT/O3/6G93lNoUd1uhz8XipZupWNWyVqw+y/QRQBEkC2T24QNfklGO7Wd5V2Msdv9KTjtPezFCunElJljBWldUeN4OFNs4wrwOXCk8hgFMWcDSprLQBqKi3Bnx4ZyXbaIZ2nCNgmOLGkJKuSq07y3iFZNgR2BeIqEQSwiwxg6i4Y9gyylqHIA1Nauiio1elneeUvgxrlLxH+/sEIsNHdpq2fKD0X8IrUkIekr5RM6wyd0RYOIMYJAsqmLUo6KFpgF6BJj6GVx748E9cR5OLy5i856R/oDFkDEfsyx6hGW6gL7vGbarxrVsRWi71KA5+9iOGXxTF7I/xPQsQ21dLOQoM6ibQ+7rssPqx58nzyrtdvAu/RXR6CFHxCUOvHQWDg== X-Forefront-Antispam-Report: CIP:216.228.117.160; CTRY:US; LANG:en; SCL:1; SRV:; IPV:NLI; SFV:NSPM; H:mail.nvidia.com; PTR:dc6edge1.nvidia.com; CAT:NONE; SFS:(13230040)(376014)(82310400026)(36860700016)(1800799024)(23010399003)(5023799004)(11063799006)(56012099006)(3023799007)(18002099003)(6133799003)(13003099007); DIR:OUT; SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: mJr7Oj/Y08doI7IctCqg5l/mF2B2C5aPPdfwDIAxlLpS25lnha/I1/sHNlZWOaLkEa/lzCc26yb8nIh85irfI5bJiKnNq9XhXw8f3NTt8lpsIBi7qIhjx3Ocs9v+srPA2Zz2S5PtZM+CM0JUSjP2hzzS6uuNvZM4X0ldFr4CnZro1WT4AzoVDpYm/PsbnR7S3vB5bvfkNHLYcKJSh87SKM30WgLHrjY9IOx9VO92/tJZxFwpxT3tbXfIjzLqB6ejVr/RyjRqspn6Yuj1mzcFg6GLt06ojIyMV+zu0dikhsuNCVxNyl17nu5qK/aFQE+KWBMhwlQIJk2oECmaVstaBs8DSMU2Gz+PLtJNqfKHbS+u/finB3ggjb8VekzXZqeIcHC9haubBimZw4gANx/TAnLbHta1k9GPdST/CIo/2Y83Cel1iq5bVuGrpAnuNm0m X-OriginatorOrg: Nvidia.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 08 Jul 2026 06:05:12.6341 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: b8559cb8-10f7-4a87-bf03-08dedcb6dfa9 X-MS-Exchange-CrossTenant-Id: 43083d15-7273-40c1-b7db-39efd9ccc17a X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=43083d15-7273-40c1-b7db-39efd9ccc17a; Ip=[216.228.117.160]; Helo=[mail.nvidia.com] X-MS-Exchange-CrossTenant-AuthSource: DS1PEPF00017095.namprd03.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: SJ1PR12MB6100 X-Spam-Status: No, score=-8.0 required=5.0 tests=BAYES_00, DKIMWL_WL_HIGH, DKIM_SIGNED, DKIM_VALID, DKIM_VALID_AU, DKIM_VALID_EF, FORGED_SPF_HELO, GIT_PATCH_0, LOCAL_AUTHENTICATION_FAIL_SPF, RCVD_IN_DNSWL_NONE, RCVD_IN_MSPIKE_H2, SPF_HELO_PASS, SPF_NONE, TXREP shortcircuit=no autolearn=ham autolearn_force=no version=3.4.6 X-Spam-Checker-Version: SpamAssassin 3.4.6 (2021-04-09) on sourceware.org X-BeenThere: gcc-patches@gcc.gnu.org X-Mailman-Version: 2.1.30 Precedence: list List-Id: Gcc-patches mailing list List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: gcc-patches-bounces~patchwork=sourceware.org@gcc.gnu.org From: Kyrylo Tkachov A comparison mask m = cmp ? CST1 : CST2 feeding several tests if (m == C) is kept as a constant-pair PHI, so each test is a separate diamond and the backward jump threader tail-duplicates the chain; the selects then come out as branches instead of conditional moves (PR125672). Let phiopt_early_allow accept the CMP, (convert)(compare) and -(convert)(compare) forms: sequences consisting only of conversions, negations and comparisons. An equality test of the materialized mask folds to the comparison itself, exposing the shared condition for the pass_merge_diamonds change that follows. match_simplify_replacement gains two checks for the new forms: * the user predictor check now also covers the multi-statement mask sequences; * a mask is not materialized for a diamond that other paths enter, through its join or through a shared arm. That would leave a PHI mixing the mask with the other paths' values, which nothing downstream recovers, and lose the ccmp chains tree-ifcombine forms from the original branches. Testsuite: pr21090.c, pr68198.c, predict-15.c and pr109071_6.c keep testing their original shapes with phiopt1 disabled, and new pr21090-2.c covers the improved folding; vrp102.c scans evrp, where the tested property is now visible; auto-init-uninit-15.c now warns at the real use instead of the inlined location, so its previously xfailed alternative becomes the expectation. Run-time neutral on 731.astcenc_r by itself; it enables the 20% improvement delivered by the following patch. Bootstrapped and regression tested on aarch64-unknown-linux-gnu. gcc/ChangeLog: PR tree-optimization/125672 PR tree-optimization/126043 * tree-ssa-phiopt.cc (phiopt_early_allow): Accept CMP, (convert)(compare) and -(convert)(compare) sequences consisting only of conversions, negations and comparisons. (match_simplify_replacement): Extend the predictor keeping check to the new mask forms. Reject materializing a mask for a diamond whose join has additional predecessors or whose arm is shared with another path. gcc/testsuite/ChangeLog: PR tree-optimization/125672 PR tree-optimization/126043 * gcc.dg/tree-ssa/pr125672-1.c: New test. * gcc.dg/tree-ssa/pr21090-2.c: New test. * gcc.dg/tree-ssa/pr21090.c: Disable phiopt1. * gcc.dg/tree-ssa/pr68198.c: Likewise. * gcc.dg/predict-15.c: Likewise. * gcc.dg/pr109071_6.c: Likewise. * gcc.dg/tree-ssa/vrp102.c: Scan evrp instead of vrp1. * gcc.dg/auto-init-uninit-15.c: Expect the warning at the use in bar instead of the inlined location. Signed-off-by: Kyrylo Tkachov --- gcc/testsuite/gcc.dg/auto-init-uninit-15.c | 12 ++-- gcc/testsuite/gcc.dg/pr109071_6.c | 5 +- gcc/testsuite/gcc.dg/predict-15.c | 4 +- gcc/testsuite/gcc.dg/tree-ssa/pr125672-1.c | 42 ++++++++++++ gcc/testsuite/gcc.dg/tree-ssa/pr21090-2.c | 27 ++++++++ gcc/testsuite/gcc.dg/tree-ssa/pr21090.c | 5 +- gcc/testsuite/gcc.dg/tree-ssa/pr68198.c | 4 +- gcc/testsuite/gcc.dg/tree-ssa/vrp102.c | 4 +- gcc/tree-ssa-phiopt.cc | 80 ++++++++++++++++++---- 9 files changed, 158 insertions(+), 25 deletions(-) create mode 100644 gcc/testsuite/gcc.dg/tree-ssa/pr125672-1.c create mode 100644 gcc/testsuite/gcc.dg/tree-ssa/pr21090-2.c diff --git a/gcc/testsuite/gcc.dg/auto-init-uninit-15.c b/gcc/testsuite/gcc.dg/auto-init-uninit-15.c index 121f0cff274..8d5fdc233b3 100644 --- a/gcc/testsuite/gcc.dg/auto-init-uninit-15.c +++ b/gcc/testsuite/gcc.dg/auto-init-uninit-15.c @@ -1,16 +1,14 @@ /* PR tree-optimization/17506 - We issue an uninitialized variable warning at a wrong location at - line 11, which is very confusing. Make sure we print out a note to - make it less confusing. (not xfailed alternative) - But it is of course ok if we warn in bar about uninitialized use - of j. (not xfailed alternative) */ + We used to issue an uninitialized variable warning at a wrong (inlined) + location, which was very confusing. Since early phiopt materializes + foo's comparison the warning is issued at the real use in bar, with the note pointing at the declaration. */ /* { dg-do compile } */ /* { dg-options "-O1 -Wuninitialized -ftrivial-auto-var-init=zero" } */ inline int foo (int i) { - if (i) /* { dg-warning "used uninitialized" } */ + if (i) return 1; return 0; } @@ -21,6 +19,6 @@ void bar (void) { int j; /* { dg-message "note: 'j' was declared here" "" } */ - for (; foo (j); ++j) /* { dg-warning "'j' is used uninitialized" "" { xfail *-*-* } } */ + for (; foo (j); ++j) /* { dg-warning "'j' is used uninitialized" } */ baz (); } diff --git a/gcc/testsuite/gcc.dg/pr109071_6.c b/gcc/testsuite/gcc.dg/pr109071_6.c index eddf15b350c..6062882583f 100644 --- a/gcc/testsuite/gcc.dg/pr109071_6.c +++ b/gcc/testsuite/gcc.dg/pr109071_6.c @@ -1,7 +1,10 @@ /* PR tree-optimization/109071 need more context for -Warray-bounds warnings due to code duplication from jump threading. test case is from PR117179, which is a duplication of PR109071. */ -/* { dg-options "-O2 -Warray-bounds -fdiagnostics-show-context=1" } */ +/* { dg-options "-O2 -Warray-bounds -fdiagnostics-show-context=1 -fdisable-tree-phiopt1" } */ +/* phiopt1 is disabled: early phiopt now materializes setval_internal's + comparison, so no threaded path with the out-of-bounds range reaches + the array access and no warning is issued at all. */ /* { dg-additional-options "-fdiagnostics-show-line-numbers -fdiagnostics-path-format=inline-events -fdiagnostics-show-caret" } */ /* { dg-enable-nn-line-numbers "" } */ const char* commands[] = {"a", "b"}; diff --git a/gcc/testsuite/gcc.dg/predict-15.c b/gcc/testsuite/gcc.dg/predict-15.c index 2a8c3ea8597..9a25537c9b6 100644 --- a/gcc/testsuite/gcc.dg/predict-15.c +++ b/gcc/testsuite/gcc.dg/predict-15.c @@ -1,5 +1,7 @@ /* { dg-do compile } */ -/* { dg-options "-O2 -fdump-tree-profile_estimate" } */ +/* { dg-options "-O2 -fdump-tree-profile_estimate -fdisable-tree-phiopt1" } */ +/* phiopt1 is disabled: early phiopt now materializes the comparison + and no goto edge remains to predict. */ int main(int argc, char **argv) { diff --git a/gcc/testsuite/gcc.dg/tree-ssa/pr125672-1.c b/gcc/testsuite/gcc.dg/tree-ssa/pr125672-1.c new file mode 100644 index 00000000000..6cc90f656c5 --- /dev/null +++ b/gcc/testsuite/gcc.dg/tree-ssa/pr125672-1.c @@ -0,0 +1,42 @@ +/* Early phiopt materializes a comparison mask "cmp ? CST1 : CST2" as + -(int)cmp / (int)cmp, exposing its comparison. An equality test of the + mask ("m == C") then folds to the comparison itself, so the mask + disappears and diamonds controlled by the same mask share an identical + condition (a prerequisite for pass_merge_diamonds). */ +/* { dg-do compile } */ +/* { dg-options "-O2 -fdump-tree-phiopt1" } */ + +/* -(int)cmp all-ones mask tested against a constant: the test folds to the + comparison. */ +int +neg_masktest (int a, int b, int x, int y) +{ + int m; + if (a > b) m = -1; else m = 0; + return (m == -1) ? x : y; +} + +/* (int)cmp boolean mask tested against a constant (the convert case). */ +int +cvt_masktest (int a, int b, int x, int y) +{ + int m; + if (a > b) m = 1; else m = 0; + return (m == 1) ? x : y; +} + +/* A mask consumed by other operators is materialized too. */ +int +mask_algebra (int a, int b) +{ + int m; + if (a > b) m = -1; else m = 0; + return m & 7; +} + +/* All three masks are materialized: no constant-pair mask PHI survives and + the equality tests fold away. */ +/* { dg-final { scan-tree-dump-not "= PHI <-1\\(" "phiopt1" } } */ +/* { dg-final { scan-tree-dump-not "== -1" "phiopt1" } } */ +/* { dg-final { scan-tree-dump-not "== 1" "phiopt1" } } */ +/* { dg-final { scan-tree-dump "& 7" "phiopt1" } } */ diff --git a/gcc/testsuite/gcc.dg/tree-ssa/pr21090-2.c b/gcc/testsuite/gcc.dg/tree-ssa/pr21090-2.c new file mode 100644 index 00000000000..d3792d4314f --- /dev/null +++ b/gcc/testsuite/gcc.dg/tree-ssa/pr21090-2.c @@ -0,0 +1,27 @@ +/* With early phiopt lowering "p != 0 ? 1 : 0" to + a boolean mask, the non-null test of a PHI of two object addresses folds + away entirely, without needing VRP's predicate folding (compare + pr21090.c). */ +/* { dg-do compile } */ +/* { dg-options "-O2 -fno-thread-jumps -fdelete-null-pointer-checks -fdump-tree-optimized" } */ + +int g, h; + +int +foo (int a) +{ + int *p; + + if (a) + p = &g; + else + p = &h; + + if (p != 0) + return 1; + else + return 0; +} + +/* { dg-final { scan-tree-dump "return 1;" "optimized" { target { ! keeps_null_pointer_checks } } } } */ +/* { dg-final { scan-tree-dump-not "PHI" "optimized" { target { ! keeps_null_pointer_checks } } } } */ diff --git a/gcc/testsuite/gcc.dg/tree-ssa/pr21090.c b/gcc/testsuite/gcc.dg/tree-ssa/pr21090.c index 92a87688601..ef755e85270 100644 --- a/gcc/testsuite/gcc.dg/tree-ssa/pr21090.c +++ b/gcc/testsuite/gcc.dg/tree-ssa/pr21090.c @@ -1,5 +1,8 @@ /* { dg-do compile } */ -/* { dg-options "-O2 -fno-thread-jumps -fdisable-tree-evrp -fdump-tree-vrp1 -fdelete-null-pointer-checks" } */ +/* { dg-options "-O2 -fno-thread-jumps -fdisable-tree-evrp -fdump-tree-vrp1 -fdelete-null-pointer-checks -fdisable-tree-phiopt1" } */ +/* phiopt1 is disabled so the non-null test reaches VRP as a branch; early + phiopt now lowers it to a boolean mask that is folded even earlier, + which pr21090-2.c covers. */ int g, h; diff --git a/gcc/testsuite/gcc.dg/tree-ssa/pr68198.c b/gcc/testsuite/gcc.dg/tree-ssa/pr68198.c index 125072941da..3b33961268f 100644 --- a/gcc/testsuite/gcc.dg/tree-ssa/pr68198.c +++ b/gcc/testsuite/gcc.dg/tree-ssa/pr68198.c @@ -1,5 +1,7 @@ /* { dg-do compile } */ -/* { dg-options "-O2 -fdump-tree-threadfull1-details -fdisable-tree-ethread" } */ +/* { dg-options "-O2 -fdump-tree-threadfull1-details -fdisable-tree-ethread -fdisable-tree-phiopt1" } */ +/* phiopt1 is disabled: early phiopt now lowers the remove test to a boolean + mask and these threading opportunities change shape. */ extern void abort (void); diff --git a/gcc/testsuite/gcc.dg/tree-ssa/vrp102.c b/gcc/testsuite/gcc.dg/tree-ssa/vrp102.c index fb62e570bed..83f6d98796f 100644 --- a/gcc/testsuite/gcc.dg/tree-ssa/vrp102.c +++ b/gcc/testsuite/gcc.dg/tree-ssa/vrp102.c @@ -1,5 +1,5 @@ /* { dg-do compile } */ -/* { dg-options "-O2 -fno-tree-dominator-opts -fdump-tree-vrp1" } */ +/* { dg-options "-O2 -fno-tree-dominator-opts -fdump-tree-evrp" } */ int f(int x, int y) { @@ -15,4 +15,4 @@ int f(int x, int y) /* We should have computed x ^ y as zero and propagated the result into the PHI feeding the result. */ -/* { dg-final { scan-tree-dump "ret_\[0-9\]+ = PHI <\[01\]\\\(\[0-9\]+\\\), \[01\]\\\(\[0-9\]+\\\)>" "vrp1" } } */ +/* { dg-final { scan-tree-dump "ret_\[0-9\]+ = PHI <\[01\]\\\(\[0-9\]+\\\), \[01\]\\\(\[0-9\]+\\\)>" "evrp" } } */ diff --git a/gcc/tree-ssa-phiopt.cc b/gcc/tree-ssa-phiopt.cc index e12dc7a8b0c..df658b04cfb 100644 --- a/gcc/tree-ssa-phiopt.cc +++ b/gcc/tree-ssa-phiopt.cc @@ -666,6 +666,29 @@ phiopt_early_allow (gimple_seq &seq, gimple_match_op &op) code = gimple_assign_rhs_code (stmt); return code == MIN_EXPR || code == MAX_EXPR; } + /* Accept CMP, (convert)(compare) and -(convert)(compare) when the + generated stmts are only conversions and comparisons, so a mask such as + "cmp ? -1 : 0" / "cmp ? 1 : 0" can be materialized as -(int)cmp / + (int)cmp, exposing its comparison. This is only the structural test; + match_simplify_replacement additionally keeps user predictors alive + and leaves diamonds that other paths enter alone. */ + if (code == NEGATE_EXPR + || CONVERT_EXPR_CODE_P (code) + || TREE_CODE_CLASS (code) == tcc_comparison) + { + for (gimple_stmt_iterator gsi = gsi_start (seq); !gsi_end_p (gsi); + gsi_next (&gsi)) + { + gimple *s = gsi_stmt (gsi); + if (!is_gimple_assign (s)) + return false; + tree_code scode = gimple_assign_rhs_code (s); + if (!CONVERT_EXPR_CODE_P (scode) + && TREE_CODE_CLASS (scode) != tcc_comparison) + return false; + } + return true; + } return false; } @@ -684,7 +707,8 @@ phiopt_early_allow (gimple_seq &seq, gimple_match_op &op) case FIXED_CST: return true; default: - return false; + /* A bare comparison (CMP) is also allowed. */ + return TREE_CODE_CLASS (code) == tcc_comparison; } } @@ -1101,9 +1125,9 @@ match_simplify_replacement (basic_block cond_bb, basic_block middle_bb, } /* For early phiopt, we don't want to lose user generated predictors - if the phiopt is converting `if (a)` into `a` as that might - be jump threaded later on so we want to keep around the - predictors. */ + if the phiopt is converting `if (a)` into `a` (or into the new + (convert)cmp / -(convert)cmp mask forms) as that might be jump + threaded later on so we want to keep around the predictors. */ if (early_p && result && TREE_CODE (result) == SSA_NAME) { bool check_it = false; @@ -1111,15 +1135,33 @@ match_simplify_replacement (basic_block cond_bb, basic_block middle_bb, tree cmp1 = gimple_cond_rhs (stmt); if (result == cmp0 || result == cmp1) check_it = true; - else if (gimple_seq_singleton_p (seq)) - { - gimple *stmt = gimple_seq_first_stmt (seq); - if (is_gimple_assign (stmt) - && result == gimple_assign_lhs (stmt) - && TREE_CODE_CLASS (gimple_assign_rhs_code (stmt)) - == tcc_comparison) + else if (!gimple_seq_empty_p (seq)) + { + /* A comparison mask: the generated statements are only + conversions, negations and comparisons, with at least one + comparison (the forms phiopt_early_allow accepts). */ + bool has_cmp = false, only_mask_ops = true; + for (gimple_stmt_iterator gi = gsi_start (seq); !gsi_end_p (gi); + gsi_next (&gi)) + { + gimple *gs = gsi_stmt (gi); + if (!is_gimple_assign (gs)) + { + only_mask_ops = false; + break; + } + tree_code c = gimple_assign_rhs_code (gs); + if (TREE_CODE_CLASS (c) == tcc_comparison) + has_cmp = true; + else if (!CONVERT_EXPR_CODE_P (c) && c != NEGATE_EXPR) + { + only_mask_ops = false; + break; + } + } + if (only_mask_ops && has_cmp) check_it = true; - } + } if (!check_it) ; else if (contains_hot_cold_predict (middle_bb)) @@ -1128,6 +1170,20 @@ match_simplify_replacement (basic_block cond_bb, basic_block middle_bb, && middle_bb != middle_bb_alt && contains_hot_cold_predict (middle_bb_alt)) return false; + /* Materializing a comparison mask for a diamond that other paths + enter (through its join or through a shared arm) would leave + a PHI mixing the mask with the other paths' values: a residual + half-diamond that neither ifcombine nor RTL if-conversion + recovers, whereas they do handle the original nested branches + (e.g. into ccmp chains). Leave such embedded diamonds alone. */ + else if (result != cmp0 && result != cmp1 + && (EDGE_COUNT (gimple_bb (phi)->preds) != 2 + || (middle_bb != gimple_bb (phi) + && !single_pred_p (middle_bb)) + || (threeway_p + && middle_bb_alt != gimple_bb (phi) + && !single_pred_p (middle_bb_alt)))) + return false; } if (!result)