Tests of General Relativity with the Binary Black Hole Signals from the LIGO-Virgo Catalog GWTC-1

The LIGO Scientific Collaboration, the Virgo Collaboration, B. P. Abbott, R. Abbott, T. D. Abbott, S. Abraham, F. Acernese, K. Ackley, C. Adams, R. X. Adhikari, V. B. Adya, C. Affeldt, M. Agathos, K. Agatsuma, N. Aggarwal, O. D. Aguiar, L. Aiello, A. Ain, P. Ajith, G. Allen, A. Allocca, M. A. Aloy, P. A. Altin, A. Amato, A. Ananyeva, S. B. Anderson, W. G. Anderson, S. V. Angelova, S. Antier, S. Appert, K. Arai, M. C. Araya, J. S. Areeda, M. Arène, N. Arnaud, K. G. Arun, S. Ascenzi, G. Ashton, S. M. Aston, P. Astone, F. Aubin, P. Aufmuth, K. AultONeal, C. Austin, V. Avendano, A. Avila-Alvarez, S. Babak, P. Bacon, F. Badaracco, M. K. M. Bader, S. Bae, P. T. Baker, F. Baldaccini, G. Ballardin, S. W. Ballmer, S. Banagiri, J. C. Barayoga, S. E. Barclay, B. C. Barish, D. Barker, K. Barkett, S. Barnum, F. Barone, B. Barr, L. Barsotti, M. Barsuglia, D. Barta, J. Bartlett, I. Bartos, R. Bassiri, A. Basti, M. Bawaj, J. C. Bayley, M. Bazzan, B. Bécsy, M. Bejger, I. Belahcene, A. S. Bell, D. Beniwal, B. K. Berger, G. Bergmann, S. Bernuzzi, J. J. Bero, C. P. L. Berry, D. Bersanetti, A. Bertolini, J. Betzwieser, R. Bhandare, J. Bidler, I. A. Bilenko, S. A. Bilgili, G. Billingsley, J. Birch, R. Birney, O. Birnholtz, S. Biscans, S. Biscoveanu, A. Bisht, M. Bitossi, M. A. Bizouard, J. K. Blackburn, C. D. Blair, D. G. Blair, R. M. Blair, S. Bloemen, N. Bode, M. Boer, Y. Boetzel, G. Bogaert, F. Bondu, E. Bonilla, R. Bonnand, P. Booker, B. A. Boom, C. D. Booth, R. Bork, V. Boschi, S. Bose, K. Bossie, V. Bossilkov, J. Bosveld, Y. Bouffanais, A. Bozzi, C. Bradaschia, P. R. Brady, A. Bramley, M. Branchesi, J. E. Brau, M. Breschi, T. Briant, J. H. Briggs, F. Brighenti, A. Brillet, M. Brinkmann, V. Brisson, R. Brito, P. Brockill, A. F. Brooks, D. D. Brown, S. Brunett, A. Buikema, T. Bulik, H. J. Bulten, A. Buonanno, D. Buskulic, M. J. Bustamante Rosell, C. Buy, R. L. Byer, M. Cabero, L. Cadonati, G. Cagnoli, C. Cahillane, J. Calderón Bustillo, T. A. Callister, E. Calloni, J. B. Camp, W. A. Campbell, M. Canepa, K. C. Cannon, H. Cao, J. Cao, C. D. Capano, E. Capocasa, F. Carbognani, S. Caride, M. F. Carney, G. Carullo, J. Casanueva Diaz, C. Casentini, S. Caudill, M. Cavaglià, F. Cavalier, R. Cavalieri, G. Cella, P. Cerdá-Durán, G. Cerretani, E. Cesarini, O. Chaibi, K. Chakravarti, S. J. Chamberlin, M. Chan, S. Chao, P. Charlton, E. A. Chase, E. Chassande-Mottin, D. Chatterjee, M. Chaturvedi, K. Chatziioannou, B. D. Cheeseboro, H. Y. Chen, X. Chen, Y. Chen, H. -P. Cheng, C. K. Cheong, H. Y. Chia, A. Chincarini, A. Chiummo, G. Cho, H. S. Cho, M. Cho, N. Christensen, Q. Chu, S. Chua, K. W. Chung, S. Chung, G. Ciani, A. A. Ciobanu, R. Ciolfi, F. Cipriano, A. Cirone, F. Clara, J. A. Clark, P. Clearwater, F. Cleva, C. Cocchieri, E. Coccia, P. -F. Cohadon, D. Cohen, R. Colgan, M. Colleoni, C. G. Collette, C. Collins, L. R. Cominsky, M. Constancio, L. Conti, S. J. Cooper, P. Corban, T. R. Corbitt, I. Cordero-Carrión, K. R. Corley, N. Cornish, A. Corsi, S. Cortese, C. A. Costa, R. Cotesta, M. W. Coughlin, S. B. Coughlin, J. -P. Coulon, S. T. Countryman, P. Couvares, P. B. Covas, E. E. Cowan, D. M. Coward, M. J. Cowart, D. C. Coyne, R. Coyne, J. D. E. Creighton, T. D. Creighton, J. Cripe, M. Croquette, S. G. Crowder, T. J. Cullen, A. Cumming, L. Cunningham, E. Cuoco, T. Dal Canton, G. Dálya, S. L. Danilishin, S. D'Antonio, K. Danzmann, A. Dasgupta, C. F. Da Silva Costa, L. E. H. Datrier, V. Dattilo, I. Dave, M. Davier, D. Davis, E. J. Daw, D. DeBra, M. Deenadayalan, J. Degallaix, M. De Laurentis, S. Deléglise, W. Del Pozzo, L. M. DeMarchi, N. Demos, T. Dent, R. De Pietri, J. Derby, R. De Rosa, C. De Rossi, R. DeSalvo, O. de Varona, S. Dhurandhar, M. C. Díaz, T. Dietrich, L. Di Fiore, M. Di Giovanni, T. Di Girolamo, A. Di Lieto, B. Ding, S. Di Pace, I. Di Palma, F. Di Renzo, A. Dmitriev, Z. Doctor, F. Donovan, K. L. Dooley, S. Doravari, I. Dorrington, T. P. Downes, M. Drago, J. C. Driggers, Z. Du, J. -G. Ducoin, P. Dupej, S. E. Dwyer, P. J. Easter, T. B. Edo, M. C. Edwards, A. Effler, P. Ehrens, J. Eichholz, S. S. Eikenberry, M. Eisenmann, R. A. Eisenstein, R. C. Essick, H. Estelles, D. Estevez, Z. B. Etienne, T. Etzel, M. Evans, T. M. Evans, V. Fafone, H. Fair, S. Fairhurst, X. Fan, S. Farinon, B. Farr, W. M. Farr, E. J. Fauchon-Jones, M. Favata, M. Fays, M. Fazio, C. Fee, J. Feicht, M. M. Fejer, F. Feng, A. Fernandez-Galiana, I. Ferrante, E. C. Ferreira, T. A. Ferreira, F. Ferrini, F. Fidecaro, I. Fiori, D. Fiorucci, M. Fishbach, R. P. Fisher, J. M. Fishner, M. Fitz-Axen, R. Flaminio, M. Fletcher, E. Flynn, H. Fong, J. A. Font, P. W. F. Forsyth, J. -D. Fournier, S. Frasca, F. Frasconi, Z. Frei, A. Freise, R. Frey, V. Frey, P. Fritschel, V. V. Frolov, P. Fulda, M. Fyffe, H. A. Gabbard, B. U. Gadre, S. M. Gaebel, J. R. Gair, L. Gammaitoni, M. R. Ganija, S. G. Gaonkar, A. Garcia, C. García-Quirós, F. Garufi, B. Gateley, S. Gaudio, G. Gaur, V. Gayathri, G. Gemme, E. Genin, A. Gennai, D. George, J. George, L. Gergely, V. Germain, S. Ghonge, Abhirup Ghosh, Archisman Ghosh, S. Ghosh, B. Giacomazzo, J. A. Giaime, K. D. Giardina, A. Giazotto, K. Gill, G. Giordano, L. Glover, P. Godwin, E. Goetz, R. Goetz, B. Goncharov, G. González, J. M. Gonzalez Castro, A. Gopakumar, M. L. Gorodetsky, S. E. Gossan, M. Gosselin, R. Gouaty, A. Grado, C. Graef, M. Granata, A. Grant, S. Gras, P. Grassia, C. Gray, R. Gray, G. Greco, A. C. Green, R. Green, E. M. Gretarsson, P. Groot, H. Grote, S. Grunewald, P. Gruning, G. M. Guidi, H. K. Gulati, Y. Guo, A. Gupta, M. K. Gupta, E. K. Gustafson, R. Gustafson, L. Haegel, O. Halim, B. R. Hall, E. D. Hall, E. Z. Hamilton, G. Hammond, M. Haney, M. M. Hanke, J. Hanks, C. Hanna, M. D. Hannam, O. A. Hannuksela, J. Hanson, T. Hardwick, K. Haris, J. Harms, G. M. Harry, I. W. Harry, C. -J. Haster, K. Haughian, F. J. Hayes, J. Healy, A. Heidmann, M. C. Heintze, H. Heitmann, P. Hello, G. Hemming, M. Hendry, I. S. Heng, J. Hennig, A. W. Heptonstall, Francisco Hernandez Vivanco, M. Heurs, S. Hild, T. Hinderer, D. Hoak, S. Hochheim, D. Hofman, A. M. Holgado, N. A. Holland, K. Holt, D. E. Holz, P. Hopkins, C. Horst, J. Hough, E. J. Howell, C. G. Hoy, A. Hreibi, E. A. Huerta, D. Huet, B. Hughey, M. Hulko, S. Husa, S. H. Huttner, T. Huynh-Dinh, B. Idzkowski, A. Iess, C. Ingram, R. Inta, G. Intini, B. Irwin, H. N. Isa, J. -M. Isac, M. Isi, B. R. Iyer, K. Izumi, T. Jacqmin, S. J. Jadhav, K. Jani, N. N. Janthalur, P. Jaranowski, A. C. Jenkins, J. Jiang, D. S. Johnson, N. K. Johnson-McDaniel, A. W. Jones, D. I. Jones, R. Jones, R. J. G. Jonker, L. Ju, J. Junker, C. V. Kalaghatgi, V. Kalogera, B. Kamai, S. Kandhasamy, G. Kang, J. B. Kanner, S. J. Kapadia, S. Karki, K. S. Karvinen, R. Kashyap, M. Kasprzack, S. Katsanevas, E. Katsavounidis, W. Katzman, S. Kaufer, K. Kawabe, N. V. Keerthana, F. Kéfélian, D. Keitel, R. Kennedy, J. S. Key, F. Y. Khalili, H. Khan, I. Khan, S. Khan, Z. Khan, E. A. Khazanov, M. Khursheed, N. Kijbunchoo, Chunglee Kim, J. C. Kim, K. Kim, W. Kim, W. S. Kim, Y. -M. Kim, C. Kimball, E. J. King, P. J. King, M. Kinley-Hanlon, R. Kirchhoff, J. S. Kissel, L. Kleybolte, J. H. Klika, S. Klimenko, T. D. Knowles, P. Koch, S. M. Koehlenbeck, G. Koekoek, S. Koley, V. Kondrashov, A. Kontos, N. Koper, M. Korobko, W. Z. Korth, I. Kowalska, D. B. Kozak, V. Kringel, N. Krishnendu, A. Królak, G. Kuehn, A. Kumar, P. Kumar, R. Kumar, S. Kumar, L. Kuo, A. Kutynia, S. Kwang, B. D. Lackey, K. H. Lai, T. L. Lam, M. Landry, B. B. Lane, R. N. Lang, J. Lange, B. Lantz, R. K. Lanza, A. Lartaux-Vollard, P. D. Lasky, M. Laxen, A. Lazzarini, C. Lazzaro, P. Leaci, S. Leavey, Y. K. Lecoeuche, C. H. Lee, H. K. Lee, H. M. Lee, H. W. Lee, J. Lee, K. Lee, J. Lehmann, A. Lenon, N. Leroy, N. Letendre, Y. Levin, J. Li, K. J. L. Li, T. G. F. Li, X. Li, F. Lin, F. Linde, S. D. Linker, T. B. Littenberg, J. Liu, X. Liu, R. K. L. Lo, N. A. Lockerbie, L. T. London, A. Longo, M. Lorenzini, V. Loriette, M. Lormand, G. Losurdo, J. D. Lough, C. O. Lousto, G. Lovelace, M. E. Lower, H. Lück, D. Lumaca, A. P. Lundgren, R. Lynch, Y. Ma, R. Macas, S. Macfoy, M. MacInnis, D. M. Macleod, A. Macquet, F. Magaña-Sandoval, L. Magaña Zertuche, R. M. Magee, E. Majorana, I. Maksimovic, A. Malik, N. Man, V. Mandic, V. Mangano, G. L. Mansell, M. Manske, M. Mantovani, F. Marchesoni, F. Marion, S. Márka, Z. Márka, C. Markakis, A. S. Markosyan, A. Markowitz, E. Maros, A. Marquina, S. Marsat, F. Martelli, I. W. Martin, R. M. Martin, D. V. Martynov, K. Mason, E. Massera, A. Masserot, T. J. Massinger, M. Masso-Reid, S. Mastrogiovanni, A. Matas, F. Matichard, L. Matone, N. Mavalvala, N. Mazumder, J. J. McCann, R. McCarthy, D. E. McClelland, S. McCormick, L. McCuller, S. C. McGuire, J. McIver, D. J. McManus, T. McRae, S. T. McWilliams, D. Meacher, G. D. Meadors, M. Mehmet, A. K. Mehta, J. Meidam, A. Melatos, G. Mendell, R. A. Mercer, L. Mereni, E. L. Merilh, M. Merzougui, S. Meshkov, C. Messenger, C. Messick, R. Metzdorff, P. M. Meyers, H. Miao, C. Michel, H. Middleton, E. E. Mikhailov, L. Milano, A. L. Miller, A. Miller, M. Millhouse, J. C. Mills, M. C. Milovich-Goff, O. Minazzoli, Y. Minenkov, A. Mishkin, C. Mishra, T. Mistry, S. Mitra, V. P. Mitrofanov, G. Mitselmakher, R. Mittleman, G. Mo, D. Moffa, K. Mogushi, S. R. P. Mohapatra, M. Montani, C. J. Moore, D. Moraru, G. Moreno, S. Morisaki, B. Mours, C. M. Mow-Lowry, Arunava Mukherjee, D. Mukherjee, S. Mukherjee, N. Mukund, A. Mullavey, J. Munch, E. A. Muñiz, M. Muratore, P. G. Murray, A. Nagar, I. Nardecchia, L. Naticchioni, R. K. Nayak, J. Neilson, G. Nelemans, T. J. N. Nelson, M. Nery, A. Neunzert, K. Y. Ng, S. Ng, P. Nguyen, D. Nichols, A. B. Nielsen, S. Nissanke, A. Nitz, F. Nocera, C. North, L. K. Nuttall, M. Obergaulinger, J. Oberling, B. D. O'Brien, G. D. O'Dea, G. H. Ogin, J. J. Oh, S. H. Oh, F. Ohme, H. Ohta, M. A. Okada, M. Oliver, P. Oppermann, Richard J. Oram, B. O'Reilly, R. G. Ormiston, L. F. Ortega, R. O'Shaughnessy, S. Ossokine, D. J. Ottaway, H. Overmier, B. J. Owen, A. E. Pace, G. Pagano, M. A. Page, A. Pai, S. A. Pai, J. R. Palamos, O. Palashov, C. Palomba, A. Pal-Singh, Huang-Wei Pan, B. Pang, P. T. H. Pang, C. Pankow, F. Pannarale, B. C. Pant, F. Paoletti, A. Paoli, A. Parida, W. Parker, D. Pascucci, A. Pasqualetti, R. Passaquieti, D. Passuello, M. Patil, B. Patricelli, B. L. Pearlstone, C. Pedersen, M. Pedraza, R. Pedurand, A. Pele, S. Penn, C. J. Perez, A. Perreca, H. P. Pfeiffer, M. Phelps, K. S. Phukon, O. J. Piccinni, M. Pichot, F. Piergiovanni, G. Pillant, L. Pinard, M. Pirello, M. Pitkin, R. Poggiani, D. Y. T. Pong, S. Ponrathnam, P. Popolizio, E. K. Porter, J. Powell, A. K. Prajapati, J. Prasad, K. Prasai, R. Prasanna, G. Pratten, T. Prestegard, S. Privitera, G. A. Prodi, L. G. Prokhorov, O. Puncken, M. Punturo, P. Puppo, M. Pürrer, H. Qi, V. Quetschke, P. J. Quinonez, E. A. Quintero, R. Quitzow-James, F. J. Raab, H. Radkins, N. Radulescu, P. Raffai, S. Raja, C. Rajan, B. Rajbhandari, M. Rakhmanov, K. E. Ramirez, A. Ramos-Buades, Javed Rana, K. Rao, P. Rapagnani, V. Raymond, M. Razzano, J. Read, T. Regimbau, L. Rei, S. Reid, D. H. Reitze, W. Ren, F. Ricci, C. J. Richardson, J. W. Richardson, P. M. Ricker, K. Riles, M. Rizzo, N. A. Robertson, R. Robie, F. Robinet, A. Rocchi, L. Rolland, J. G. Rollins, V. J. Roma, M. Romanelli, R. Romano, C. L. Romel, J. H. Romie, K. Rose, D. Rosińska, S. G. Rosofsky, M. P. Ross, S. Rowan, A. Rüdiger, P. Ruggi, G. Rutins, K. Ryan, S. Sachdev, T. Sadecki, M. Sakellariadou, L. Salconi, M. Saleem, A. Samajdar, L. Sammut, E. J. Sanchez, L. E. Sanchez, N. Sanchis-Gual, V. Sandberg, J. R. Sanders, K. A. Santiago, N. Sarin, B. Sassolas, B. S. Sathyaprakash, P. R. Saulson, O. Sauter, R. L. Savage, P. Schale, M. Scheel, J. Scheuer, P. Schmidt, R. Schnabel, R. M. S. Schofield, A. Schönbeck, E. Schreiber, B. W. Schulte, B. F. Schutz, S. G. Schwalbe, J. Scott, S. M. Scott, E. Seidel, D. Sellers, A. S. Sengupta, N. Sennett, D. Sentenac, V. Sequino, A. Sergeev, Y. Setyawati, D. A. Shaddock, T. Shaffer, M. S. Shahriar, M. B. Shaner, L. Shao, P. Sharma, P. Shawhan, H. Shen, R. Shink, D. H. Shoemaker, D. M. Shoemaker, S. ShyamSundar, K. Siellez, M. Sieniawska, D. Sigg, A. D. Silva, L. P. Singer, N. Singh, A. Singhal, A. M. Sintes, S. Sitmukhambetov, V. Skliris, B. J. J. Slagmolen, T. J. Slaven-Blair, J. R. Smith, R. J. E. Smith, S. Somala, E. J. Son, B. Sorazu, F. Sorrentino, T. Souradeep, E. Sowell, A. P. Spencer, A. K. Srivastava, V. Srivastava, K. Staats, C. Stachie, M. Standke, D. A. Steer, M. Steinke, J. Steinlechner, S. Steinlechner, D. Steinmeyer, S. P. Stevenson, D. Stocks, R. Stone, D. J. Stops, K. A. Strain, G. Stratta, S. E. Strigin, A. Strunk, R. Sturani, A. L. Stuver, V. Sudhir, T. Z. Summerscales, L. Sun, S. Sunil, J. Suresh, P. J. Sutton, B. L. Swinkels, M. J. Szczepańczyk, M. Tacca, S. C. Tait, C. Talbot, D. Talukder, D. B. Tanner, M. Tápai, A. Taracchini, J. D. Tasson, R. Taylor, F. Thies, M. Thomas, P. Thomas, S. R. Thondapu, K. A. Thorne, E. Thrane, Shubhanshu Tiwari, Srishti Tiwari, V. Tiwari, K. Toland, M. Tonelli, Z. Tornasi, A. Torres-Forné, C. I. Torrie, D. Töyrä, F. Travasso, G. Traylor, M. C. Tringali, A. Trovato, L. Trozzo, R. Trudeau, K. W. Tsang, M. Tse, R. Tso, L. Tsukada, D. Tsuna, D. Tuyenbayev, K. Ueno, D. Ugolini, C. S. Unnikrishnan, A. L. Urban, S. A. Usman, H. Vahlbruch, G. Vajente, G. Valdes, N. van Bakel, M. van Beuzekom, J. F. J. van den Brand, C. Van Den Broeck, D. C. Vander-Hyde, J. V. van Heijningen, L. van der Schaaf, A. A. van Veggel, M. Vardaro, V. Varma, S. Vass, M. Vasúth, A. Vecchio, G. Vedovato, J. Veitch, P. J. Veitch, K. Venkateswara, G. Venugopalan, D. Verkindt, F. Vetrano, A. Viceré, A. D. Viets, D. J. Vine, J. -Y. Vinet, S. Vitale, T. Vo, H. Vocca, C. Vorvick, S. P. Vyatchanin, A. R. Wade, L. E. Wade, M. Wade, R. M. Wald, R. Walet, M. Walker, L. Wallace, S. Walsh, G. Wang, H. Wang, J. Z. Wang, W. H. Wang, Y. F. Wang, R. L. Ward, Z. A. Warden, J. Warner, M. Was, J. Watchi, B. Weaver, L. -W. Wei, M. Weinert, A. J. Weinstein, R. Weiss, F. Wellmann, L. Wen, E. K. Wessel, P. Weßels, J. W. Westhouse, K. Wette, J. T. Whelan, B. F. Whiting, C. Whittle, D. M. Wilken, D. Williams, A. R. Williamson, J. L. Willis, B. Willke, M. H. Wimmer, W. Winkler, C. C. Wipf, H. Wittel, G. Woan, J. Woehler, J. K. Wofford, J. Worden, J. L. Wright, D. S. Wu, D. M. Wysocki, L. Xiao, H. Yamamoto, C. C. Yancey, L. Yang, M. J. Yap, M. Yazback, D. W. Yeeles, Hang Yu, Haocun Yu, S. H. R. Yuen, M. Yvert, A. K. Zadrożny, M. Zanolin, T. Zelenova, J. -P. Zendri, M. Zevin, J. Zhang, L. Zhang, T. Zhang, C. Zhao, M. Zhou, Z. Zhou, X. J. Zhu, A. B. Zimmerman, M. E. Zucker, J. Zweizig

I Introduction

Einstein’s theory of gravity, general relativity (GR), has withstood a large number of experimental tests Will 2014. With the advent of gravitational-wave (GW) astronomy and the observations by the Advanced LIGO Aasi et al. 2015 and Advanced Virgo Acernese et al. 2015 detectors, a range of new tests of GR have become possible. These include both weak field tests of the propagation of GWs, as well as tests of the strong field regime of compact binary sources. See Abbott et al. 2016; Abbott et al. 2016; Abbott et al. 2017; Abbott et al. 2017; Abbott et al. 2019 for previous applications of such tests to GW data.

We report results from tests of GR on all the confident binary black hole GW events in the catalog GWTC-1 LIGO Scientific Collaboration and Virgo Collaboration 2018, i.e., those from the first and second observing runs of the advanced generation of detectors. Besides all of the events previously announced (GW150914, GW151012, GW151226, GW170104, GW170608, and GW170814) Abbott et al. 2016; Abbott et al. 2016; Abbott et al. 2016; Abbott et al. 2016; Abbott et al. 2017; Abbott et al. 2017; Abbott et al. 2017, this includes the four new GW events reported in Abbott et al. 2019 (GW170729, GW170809, GW170818, and GW170823). We do not investigate any of the marginal triggers in GWTC-1, which have a false-alarm rate (FAR) greater than one per year. Table 1 displays a complete list of the events we consider. Tests of GR on the binary neutron star event GW170817 are described in Abbott et al. 2019.

At present, there are no complete theories of gravity other than GR that are mathematically and physically viable and provide well defined alternative predictions for the waveforms arising from the coalescence of two black holes (if, indeed, these theories even admit black holes). There are very preliminary simulations of scalar waveforms from binary black holes in the effective field theory (EFT) framework in alternative theories Okounkova et al. 2017; Witek et al. 2019, and the leading corrections to the gravitational waveforms in head-on collisions Okounkova et al. 2019, but these simulations require much more development before their results can be used in gravitational wave data analysis. There are also concerns about the mathematical viability of the theories considered when they are not treated in the EFT framework. Thus, we cannot test GR by direct comparison with other specific theories. Instead, we can

These methods are agnostic to any particular choice of alternative theory. For the most part, our results should therefore be interpreted as observational constraints on possible GW phenomenologies, independent of the overall suitability or well-posedness of any specific alternative to GR. These limits are useful in providing a quantitative indication of the degree to which the data is described by GR; they may also be interpreted more specifically in the context of any given alternative to produce constraints, if applicable.

In particular, with regard to the consistency of the GR predictions (i), we

With regard to deviations from GR (ii), we separately introduce parametrizations for

The former could be viewed as representing possible GR modifications in the strong-field region close to the binary, while the latter would correspond to weak-field modifications away from the source. Although we consider these independently, modifications to GW propagation would most likely be accompanied by modifications to GW generation in any given extension of GR. We have also checked that none of the events discussed here provide stronger constraints on models with purely vector and purely scalar GW polarizations than those previously published in Abbott et al. 2017; Abbott et al. 2019. Our analyses do not reveal any inconsistency of the data with the predictions of GR. These results supersede all our previous testing GR results on the binary black hole signals found in O1 and O2 Abbott et al. 2016; Abbott et al. 2016; Abbott et al. 2017; Abbott et al. 2017. In particular, the previously published residuals and propagation test results were affected by a slight normalization issue.

Limits on deviations from GR for individual events are dominated by statistical errors due to detector noise. These errors can be reduced by appropriately combining results from multiple events. Sources of systematic errors, on the other hand, include uncertainties in the detector calibration and power spectral density (PSD) estimation and errors in the modeling of waveforms in GR. Detector calibration uncertainties are modeled as corrections to the measured detector response function and are marginalized over. Studies on the effect of PSD uncertainties are currently ongoing. A full characterization of the systematic errors due to the GR waveform models that we employ is beyond the scope of this study; some investigations can be found in Bohé et al. 2017; Khan et al. 2016; Blackman et al. 2017; Khan et al. 2019; Abbott et al. 2017.

This paper is organized as follows. Section II provides an overview of the data sets employed here, while Sec. III details which GW events are used to produce the individual and combined results presented in this paper. In Sec. IV we explain the gravitational waveforms and data analysis formalisms which our tests of GR are based on, before we present the results in the following sections. Section V contains two signal consistency tests: the residuals test in V.1 and the inspiral-merger-ringdown consistency test in V.2. Results from parameterized tests are given in Sec. VI for GW generation, and in Sec. VII for GW propagation. We briefly discuss the study of GW polarizations in Sec. VIII. Finally, we conclude in Sec. IX. We give results for individual events and some checks on waveform systematics in the Appendix.

The results of each test and associated data products can be found in Ref. LIGO Scientific Collaboration and Virgo Collaboration 2019. The GW strain data for all the events considered are available at the Gravitational Wave Open Science Center LIGO Scientific Collaboration, Virgo Collaboration 2018.

II Data, calibration and cleaning

The first observing run of Advanced LIGO (O1) lasted from September 12th, 2015 to January 19th, 2016. The second observing run (O2) lasted from November 30th, 2016 to August 25th, 2017, with the Advanced Virgo observatory joining on August 1st, 2017. This paper includes all GW events originating from the coalescence of two black holes found in these two data sets and published in Abbott et al. 2016; Abbott et al. 2019.

The GW detector’s response to changes in the differential arm length (the interferometer’s degree of freedom most sensitive to GWs) must be calibrated using independent, accurate, absolute references. The LIGO detectors use photon recoil (radiation pressure) from auxiliary laser systems to induce mirror motions that change the arm cavity lengths, allowing a direct measure of the detector response Karki et al. 2016; Cahillane et al. 2017; Viets et al. 2018. Calibration of Virgo relies on measurements of Michelson interference fringes as the main optics swing freely, using the primary laser wavelength as a fiducial length. Subsequent measurements propagate the calibration to arrive at the final detector response Estevez et al. 2018; Acernese et al. 2018. These complex-valued, frequency-dependent measurements of the LIGO and Virgo detectors’ response yield the uncertainty in their respective estimated amplitude and phase of the GW strain output. The amplitude and phase correction factors are modeled as cubic splines and marginalized over in the estimation of astrophysical source parameters Farr et al. 2015; Veitch et al. 2015; Abbott et al. 2019; Abbott et al. 2019. Additionally, the uncertainty in the time stamping of Virgo data (much larger than the LIGO timing uncertainty, which is included in the phase correction factor) is also accounted for in the analysis.

Post-processing techniques to subtract noise contributions and frequency lines from the data around gravitational-wave events were developed in O2 and introduced in Abbott et al. 2017; Abbott et al. 2017; Abbott et al. 2017, for the astrophysical parameter estimation of GW170608, GW170814, and GW170817. This noise subtraction was achieved using optimal Wiener filters to calculate coupling transfer functions from auxiliary sensors Driggers et al. 2019. A new, optimized parallelizable method in the frequency domain Davis et al. 2019 allows large scale noise subtraction on LIGO data. All of the O2 analyses presented in this manuscript use the noise-subtracted data set with the latest calibration available. The O1 data set is the same used in previous publications, as the effect of noise subtraction is expected to be negligible. Reanalysis of the O1 events is motivated by improvements in the parameter estimation pipeline, an improved frequency-dependent calibration, and the availability of new waveform models.

III Events and significance

We present results for all confident detections of binary black hole events in GTWC-1 LIGO Scientific Collaboration and Virgo Collaboration 2018, i.e., all such events detected during O1 and O2 with a FAR lower than one per year, as published in Abbott et al. 2019. The central columns of Table 1 list the FARs of each event as evaluated by the three search pipelines used in Abbott et al. 2019. Two of these pipelines (PyCBC and GstLAL) rely on waveform templates computed from binary black hole coalescences in GR. Making use of a measure of significance that assumes the validity of GR could potentially lead to biases in the selection of events to be tested, systematically disfavoring signals in which a GR violation would be most evident. Therefore, it is important to consider the possibilities that

We consider each of the GW events individually, carrying out different analyses on a case-by-case basis. Some of the tests presented here, such as the inspiral-merger-ringdown (IMR) consistency test in Sec. V.2 and the parameterized tests in Sec. VI, distinguish between the inspiral and the post-inspiral regimes of the signal. The separation between these two regimes is performed in the frequency-domain, choosing a particular cutoff frequency determined by the parameters of the event. Larger-mass systems merge at lower frequencies, presenting a short inspiral signal in band; lower mass systems have longer observable inspiral signals, but the detector’s sensitivity decreases at higher frequencies and hence the post-inspiral signal becomes less informative. Therefore, depending on the total mass of the system, a particular signal might not provide enough information within the sensitive frequency band of the GW detectors for all analyses.

As a proxy for the amount of information that can be extracted from each part of the signal, we calculate the signal-to-noise ratio (SNR) of the inspiral and the post-inspiral parts of the signals separately. We only apply inspiral (post-inspiral) tests if the inspiral (post-inspiral) SNR is greater than 6. Each test uses a different inspiral-cutoff frequency, and hence they assign different SNRs to the two regimes (details provided in the relevant section for each test). In Table 1 we indicate which analyses have been performed on which event, based on this frequency and the corresponding SNR. While we perform these tests on all events with SNR >6>6 in the appropriate regime, in a few cases the results appear uninformative and the posterior distribution extends across the entire prior considered. Since the results are prior dependent, upper limits should not be set from these individual analyses. See Sec. A.3 of the Appendix for details.

In addition to the individual analysis of each event, we derive combined constraints on departures from GR using multiple signals simultaneously. Constraints from individual events are largely dominated by statistical uncertainties due to detector noise. Combining events together can reduce such statistical errors on parameters that take consistent values across all events. However, it is impossible to make joint probabilistic statements from multiple events without prior assumptions about the nature of each observation and how it relates to others in the set. This means that, although there are well-defined statistical procedures for producing joint results, there is no unique way of doing so.

In light of this, we adopt what we take to be the most straightforward strategy, although future studies may follow different criteria. First, in combining events we assume that deviations from GR are manifested equally across events, independent of source properties. This is justified for studies of modified GW propagation, since those effects should not depend on the source. Propagation effects do depend critically on source distance. However, this dependence is factored out explicitly, in a way that allows for combining events as we do here (see Sec. VII). For other analyses, it is quite a strong assumption to take all deviations from GR to be independent of source properties. Such combined tests should not be expected to necessarily reveal generic source-dependent deviations, although they might if the departures from GR are large enough (see, e.g., Ghosh et al. 2018). Future work may circumvent this issue by combining marginalized likelihood ratios (Bayes factors), instead of posterior probability distributions Agathos et al. 2014. More general ways of combining results are discussed and implemented in Refs. Zimmerman et al. 2019; Isi et al. 2019.

Second, we choose to produce combined constraints only from events that were found in both modeled searches (PyCBC Nitz et al. 2018; Dal Canton et al. 2014; Usman et al. 2016 and GstLAL Sachdev et al. 2019; Messick et al. 2017) with a FAR of at most one per one-thousand years. This ensures that there is a very small probability of inclusion of a non-astrophysical event. The events used for the combined results are indicated with bold names in Table 1. The events thus excluded from the combined analysis have low SNR and would therefore contribute only marginally to tightened constraints. Excluding marginal events from our analyses amounts to assigning a null a priori probability to the possibility that those data contain any information about the tests in question. This is, in a sense, the most conservative choice.

IV Parameter inference

The starting point for all the analyses presented here are waveform models that describe the GWs emitted by coalescing black-hole binaries. The GW signature depends on the intrinsic parameters describing the binary as well as the extrinsic parameters specifying the location and orientation of the source with respect to the detector network. The intrinsic parameters for circularized black-hole binaries in GR are the two masses mim_{i} of the black holes and the two spin vectors S⃗i\vec{S}_{i} defining the rotation of each black hole, where i∈{1,2}i\in\{1,2\} labels the two black holes. We assume that the binary has negligible orbital eccentricity, as is expected to be the case when the binary enters the band of ground-based detectors Peters and Mathews 1963; Peters 1964 (except in some more extreme formation scenarios, These scenarios could occur often enough, compared to the expected rate of detections, that the inclusion of eccentricity in waveform models is a necessity for tests of GR in future observing runs; see, e.g., Hinderer and Babak 2017; Cao and Han 2017; Hinder et al. 2018; Huerta et al. 2018; Klein et al. 2018; Moore et al. 2018; Moore and Yunes 2019; Tiwari et al. 2019 for recent work on developing such waveform models. e.g., Samsing 2018; Rodriguez et al. 2018; Zevin et al. 2019; Rodriguez et al. 2018). The extrinsic parameters comprise four parameters that specify the space-time location of the binary black-hole, namely the sky location (right ascension and declination), the luminosity distance, and the time of coalescence. In addition, there are three extrinsic parameters that determine the orientation of the binary with respect to Earth, namely the inclination angle of the orbit with respect to the observer, the polarization angle, and the orbital phase at coalescence.

We employ two waveform families that model binary black holes in GR: the effective-one-body based SEOBNRv4 Bohé et al. 2017 waveform family that assumes non-precessing spins for the black holes (we use the frequency domain reduced order model SEOBNRv4_ROM for reasons of computational efficiency), and the phenomenological waveform family IMRPhenomPv2 Husa et al. 2016; Khan et al. 2016; Hannam et al. 2014 that models the effects of precessing spins using two effective parameters by twisting up the underlying aligned-spin model. We use IMRPhenomPv2 to obtain all the main results given in this paper, and use SEOBNRv4 to check the robustness of these results, whenever possible. When we use IMRPhenomPv2, we impose a prior m1/m2≤18m_{1}/m_{2}\leq 18 on the mass ratio, as the waveform family is not calibrated against numerical relativity simulations for m1/m2>18m_{1}/m_{2}>18. We do not impose a similar prior when using SEOBNRv4, since it includes information about the extreme mass ratio limit. Neither of these waveform models includes the full spin dynamics (which requires 6 spin parameters). Fully precessing waveform models have been recently developed Taracchini et al. 2014; Babak et al. 2017; Blackman et al. 2017; Khan et al. 2019; Varma et al. 2019 and will be used in future applications of these tests.

The waveform models used in this paper do not include the effects of subdominant (non-quadrupole) modes, which are expected to be small for comparable-mass binaries Kamaretsos et al. 2012; London et al. 2014. The first generation of binary black hole waveform models including spin and higher order modes has recently been developed Blackman et al. 2017; London et al. 2018; Cotesta et al. 2018; Varma et al. 2019; Varma et al. 2019. Preliminary results in Abbott et al. 2019, using NR simulations supplemented by NR surrogate waveforms, indicate that the higher mode content of the GW signals detected by Advanced LIGO and Virgo is weak enough that models without the effect of subdominant modes do not introduce substantial biases in the intrinsic parameters of the binary. For unequal-mass binaries, the effect of the non-quadrupole modes is more pronounced Varma and Ajith 2017, particularly when the binary’s orientation is close to edge-on. In these cases, the presence of non-quadrupole modes can show up as a deviation from GR when using waveforms that only include the quadrupole modes, as was shown in Pang et al. 2018. Applications of tests of GR with the new waveform models that include non-quadrupole modes will be carried out in the future.

We believe that our simplifying assumptions on the waveform models (zero eccentricity, simplified treatment of spins, and neglect of subdominant modes) are justified by astrophysical considerations and previous studies. Indeed, as we show in the remainder of the paper, the observed signals are consistent with the waveform models. Of course, had our analyses resulted in evidence for violations of GR, we would have had to revisit these simplifications very carefully.

The tests described in this paper are performed within the framework of Bayesian inference, by means of the LALInference code Veitch et al. 2015 in the LIGO Scientific Collaboration Algorithm Library Suite (LALSuite) LIGO Scientific Collaboration and Virgo Collaboration 2018. We estimate the PSD using the BayesWave code Cornish and Littenberg 2015; Littenberg and Cornish 2015, as described in Appendix B of Abbott et al. 2019. Except for the residuals test described in Sec. V.1, we use the waveform models described in this section to estimate from the data the posterior distributions of the parameters of the binary. These include not only the intrinsic and extrinsic parameters mentioned above, but also other parameters that describe possible departures from the GR predictions. Specifically, for the parameterized tests in Secs. VI and VII, we modify the phase Φ(f)\Phi(f) of the frequency-domain waveform

For the GR parameters, we use the same prior distributions as the main parameter estimation analysis described in Abbott et al. 2019, though for a number of the tests we need to extend the ranges of these priors to account for correlations with the non-GR parameters, or for the fact that only a portion of the signal is analyzed (as in Sec. V.2). We also use the same low-frequency cutoffs for the likelihood integral as in Abbott et al. 2019, i.e., 2020 Hz for all events except for GW170608, where 3030 Hz is used for LIGO Hanford, as discussed in Abbott et al. 2017, and GW170818, where 1616 Hz is used for all three detectors. For the model agnostic residual test described in Sec. V.1, we use the BayesWave code Cornish and Littenberg 2015 which describes the GW signals in terms of a number of Morlet-Gabor wavelets.

V Consistency tests

One way to evaluate the ability of GR to describe GW signals is to subtract the best-fit template from the data and make sure the residuals have the statistical properties expected of instrumental noise. This largely model-independent test is sensitive to a wide range of possible disagreements between the data and our waveform models, including those caused by deviations from GR and by modeling systematics. This analysis can look for GR violations without relying on specific parametrizations of the deviations, making it a versatile tool. Results from a similar study on our first detection were already presented in Abbott et al. 2016.

In order to establish whether the residuals agree with noise (Gaussian or otherwise), we proceed as follows. For each event in our set, we use LALInference and the IMRPhenomPv2 waveform family to obtain an estimate of the best-fit (i.e., maximum likelihood) binary black-hole waveform based on GR. This waveform incorporates factors that account for uncertainty in the detector calibration, as described in Sec. II. This best-fit waveform is then subtracted from the data to obtain residuals for a 1 second window centered on the trigger time reported in Abbott et al. 2019. The analysis is sensitive only to residual power in that 1 s window due to technicalities related to how BayesWave handles its sine-Gaussian basis elements Cornish and Littenberg 2015; Littenberg and Cornish 2015. If the GR-based model provides a good description of the signal, we expect the resulting residuals at each detector to lack any significant coherent SNR beyond what is expected from noise fluctuations. We compute the coherent SNR using BayesWave, which models the multi-detector residuals as a superposition of incoherent Gaussian noise and an elliptically-polarized coherent signal. The residual network SNR reported by BayesWave is the SNR that would correspond to such a coherent signal.

The set of pp-values shown in Table 2 is consistent with all coherent residual power being due to instrumental noise. Assuming that this is indeed the case, we expect the pp-values to be uniformly distributed over $,whichexplainsthevariationinTable2.Withonlytenevents,however,itisdifficulttoobtainstrongquantitativeevidenceoftheuniformityofthisdistribution.Nevertheless,wefollowFisher’smethodFisher1948tocomputeameta, which explains the variation in Table 2. With only ten events, however, it is difficult to obtain strong quantitative evidence of the uniformity of this distribution. Nevertheless, we follow Fisher’s method Fisher 1948 to compute a metap−valueforthenullhypothesisthattheindividual-value for the null hypothesis that the individualp−valuesinTable2areuniformlydistributed.Weobtainameta-values in Table 2 are uniformly distributed. We obtain a metap$-value of 0.4, implying that there is no evidence for coherent residual power that cannot be explained by noise alone. All in all, this means that there is no statistically significant evidence for deviations from GR.

V.2 Inspiral-merger-ringdown consistency test

VI Parameterized tests of gravitational wave generation

A deviation from GR could manifest itself as a modification of the dynamics of two orbiting compact objects, and in particular, the evolution of the orbital (and hence, GW) phase. In an analytical waveform model like IMRPhenomPv2, the details of the GW phase evolution are controlled by coefficients that are either analytically calculated or determined by fits to numerical-relativity (NR) simulations, always under the assumption that GR is the underlying theory. In this section we investigate deviations from the GR binary dynamics by introducing shifts in each of the individual GW phase coefficients of IMRPhenomPv2. Such shifts correspond to deviations in the waveforms from the predictions of GR. We then treat these shifts as additional unconstrained GR-violating parameters, which we measure in addition to the standard parameters describing the binary.

The early inspiral of compact binaries is well modeled by the post-Newtonian (PN) approximation Blanchet et al. 1995; Blanchet et al. 2004; Blanchet et al. 2005; Blanchet 2014 to GR, which is based on the expansion of the orbital quantities in terms of a small velocity parameter v/cv/c. For a given set of intrinsic parameters, coefficients for the different orders in v/cv/c in the PN series are uniquely determined. A consistency test of GR using measurements of the inspiral PN phase coefficients was first proposed in Arun et al. 2006; Arun et al. 2006; Mishra et al. 2010, while a generalized parametrization was motivated in Yunes and Pretorius 2009. Bayesian implementations based on such parametrized methods were presented and tested in Li et al. 2012; Li et al. 2012; Agathos et al. 2014; Cornish et al. 2011 and were also extended to the post-inspiral part of the gravitational-wave signal Sampson et al. 2014; Meidam et al. 2018. These ideas were applied to the first GW observation, GW150914 Abbott et al. 2016, yielding the first bounds on higher-order PN coefficients Abbott et al. 2016. Since then, the constraints have been revised with the binary black hole events that followed, GW151226 in O1 Abbott et al. 2016 and GW170104 in O2 Abbott et al. 2017. More recently, the first such constraints from a binary neutron star merger were placed with the detection of GW170817 Abbott et al. 2019. Bounds on parametrized violations of GR from GW detections have been mapped, to leading order, to constraints on specific alternative theories of gravity (see, e.g., Yunes et al. 2016). In this paper, we present individual constraints on parametrized deviations from GR for each of the GW sources in O1 and O2 listed in Table 1, as well as the tightest combined constraints obtained to date by combining information from all the significant binary black holes events observed so far, as described in Sec. III.

The frequency-domain GW phase evolution Φ(f)\Phi(f) in the early-inspiral stage of IMRPhenomPv2 is described by a PN expansion, augmented with higher-order phenomenological coefficients. The PN phase evolution is analytically expressed in closed form by employing the stationary phase approximation. The late-inspiral and post-inspiral (intermediate and merger-ringdown) stages are described by phenomenological analytical expressions. The transition frequency This frequency is different than the cutoff frequency used in the inspiral-merger-ringdown consistency test, as was briefly mentioned in Sec. III. from inspiral to intermediate regime is given by the condition GMf/c3=0.018GMf/c^{3}=0.018, with MM the total mass of the binary in the detector frame, since this is the lowest frequency above which this model was calibrated with NR data Khan et al. 2016. Let us use pip_{i} to collectively denote all of the inspiral and post-inspiral parameters φi\varphi_{i}, αi\alpha_{i}, βi\beta_{i}, that will be introduced below. Deviations from GR in all stages are expressed by means of relative shifts δp^i\delta\hat{p}_{i} in the corresponding waveform coefficients: pi→(1+δp^i) pip_{i}\rightarrow(1+\delta\hat{p}_{i})\ p_{i}, which are used as additional free parameters in our extended waveform models.

We denote the testing parameters corresponding to PN phase coefficients by δφ^i\delta\hat{\varphi}_{i}, where ii indicates the power of v/cv/c beyond leading (Newtonian or 00PN) order in Φ(f)\Phi(f). The frequency dependence of the corresponding phase term is f(i−5)/3f^{(i-5)/3}. In the parametrized model, ii varies from 0 to 7, including the terms with logarithmic dependence at 2.52.5PN and 33PN. The non-logarithmic term at 2.52.5PN (i.e., i=5i=5) cannot be constrained, because of its degeneracy with a constant reference phase (e.g., the phase at coalescence). These coefficients were introduced in their current form in Eq. (19) of Li et al. 2012. In addition, we also test for i=−2i=-2, representing an effective −1-1PN term, which is motivated below. The full set of inspiral parameters are thus

Since the −1-1PN term and the 0.50.5PN term are absent in the GR phasing, we parametrize δφ^−2\delta\hat{\varphi}_{-2} and δφ^1\delta\hat{\varphi}_{1} as absolute deviations, with a pre-factor equal to the 00PN coefficient.

To measure the above GR violations in the post-Newtonian inspiral, we employ two waveform models: (i) the analytical frequency-domain model IMRPhenomPv2 which also provided the natural parametrization for our tests and (ii) SEOBNRv4, which we use in the form of SEOBNRv4_ROM, a frequency-domain, reduced-order-model of the SEOBNRv4 model. The inspiral part of SEOBNRv4 is based on a numerical evolution of the aligned-spin effective-one-body dynamics of the binary and its post-inspiral model is phenomenological. The entire SEOBNRv4 model is calibrated against NR simulations. Despite its non-analytical nature, SEOBNRv4_ROM can also be used to test the parametrized modifications of the early inspiral defined above. Using the method presented in Abbott et al. 2019, we add deviations to the waveform phase corresponding to a given δφ^i\delta\hat{\varphi}_{i} at low frequencies and then taper the corrections to zero at a frequency consistent with the transition frequency between early-inspiral and intermediate phases used by IMRPhenomPv2. The same procedure cannot be applied to the later stages of the waveform, thus the analysis performed with SEOBNRv4 is restricted to the post-Newtonian inspiral, cf. Fig. 3.

The analytical descriptions of the intermediate and merger-ringdown stages in the IMRPhenomPv2 model allow for a straightforward way of parametrizing deviations from GR, denoted by {δβ^2,δβ^3}\{\delta\hat{\beta}_{2},\delta\hat{\beta}_{3}\} and {δα^2,δα^3,δα^4}\{\delta\hat{\alpha}_{2},\delta\hat{\alpha}_{3},\delta\hat{\alpha}_{4}\} respectively, following Meidam et al. 2018. Here the parameters δβ^i\delta\hat{\beta}_{i} correspond to deviations from the NR-calibrated phenomenological coefficients βi\beta_{i} of the intermediate stage, while the parameters δα^i\delta\hat{\alpha}_{i} refer to modifications of the merger-ringdown coefficients αi\alpha_{i} obtained from a combination of phenomenological fits and analytical black-hole perturbation theory calculations Khan et al. 2016.

Using LALInference, we calculate posterior distributions of the parameters characterizing the waveform (including those that describe the binary in GR). Our parametrization recovers GR at δp^i=0\delta\hat{p}_{i}=0, so consistency with GR is verified if the posteriors of δp^i\delta\hat{p}_{i} have support at zero. We perform the analyses by varying one δp^i\delta\hat{p}_{i} at a time; as shown in Ref. Sampson et al. 2013, this is fully robust to detecting deviations present in multiple PN-orders. In addition, allowing for a larger parameter space by varying multiple coefficients simultaneously would not improve our efficiency in identifying violations of GR, as it would yield less informative posteriors. A specific alternative theory of gravity would likely yield correlated deviations in many parameters, including modifications that we have not considered here. This would be the target of an exact comparison of an alternative theory with GR, which would only be possible if a complete, accurate description of the inspiral-merger-ringdown signal in that theory was available.

Figure 4 shows the 90%90\% upper bounds on ∣δφ^i∣|\delta\hat{\varphi}_{i}| for all the individual events which cross the SNR threshold (SNR >> 6) in the inspiral regime (the most massive of which is GW150914). The bounds from the combined posteriors are also shown; these include the events which exceed both the SNR threshold in the inspiral regime as well as the significance threshold, namely GW150914, GW151226, GW170104, GW170608, and GW170814. The bound from the likely lightest mass binary black hole event GW170608 at 1.5PN is currently the strongest constraint obtained on a positive PN coefficient from a single binary black hole event, as shown in Fig. 4. However, the constraint at this order is about five times worse than that obtained from the binary neutron star event GW170817 alone Abbott et al. 2019. The −1-1PN bound is two orders of magnitude better for GW170817 than the best bound obtained here (from GW170608). The corresponding best −1-1PN bound coming from the double pulsar PSR J0737−-3039, is a few orders of magnitude tighter still, at ∣δφ^−2∣≲10−7|\delta\hat{\varphi}_{-2}|\lesssim 10^{-7} Barausse et al. 2016; Kramer et al. 2006. At 00PN we find that the bound from GW170608 beats the one from GW170817, but remains weaker than the one from the double pulsar by one order of magnitude Yunes and Hughes 2010; Kramer et al. 2006. For all other PN orders, GW170608 also provides the best bounds, which at high PN orders are of the same order of magnitude as the ones from GW170817. Our results can be compared statistically to those obtained by performing the same tests on simulated GR and non-GR waveforms given in Meidam et al. 2018. The results presented here are consistent with those of GR waveforms injected into realistic detector data. The combined bounds are the tightest obtained so far, improving on the bounds obtained in Abbott et al. 2016 by factors between 1.1 and 1.8.

VII Parameterized tests of gravitational wave propagation

We now place constraints on a phenomenological modification of the GW dispersion relation, i.e., on a possible frequency dependence of the speed of GWs. This modification, introduced in Mirshekari et al. 2012 and first applied to LIGO data in Abbott et al. 2017, is obtained by adding a power-law term in the momentum to the dispersion relation E2=p2c2E^{2}=p^{2}c^{2} of GWs in GR, giving

Here, cc is the speed of light, EE and pp are the energy and momentum of the GWs, and AαA_{\alpha} and α\alpha are phenomenological parameters. We consider α\alpha values from 00 to 44 in steps of 0.50.5. However, we exclude α=2\alpha=2, where the speed of the GWs is modified in a frequency-independent manner, and therefore gives no observable dephasing. For a source with an electromagnetic counterpart, A2A_{2} can be constrained by comparison with the arrival time of the photons, as was done with GW170817/GRB170817A Abbott et al. 2017. Thus, in all cases except for α=0\alpha=0, we are considering a Lorentz-violating dispersion relation. The group velocity associated with this dispersion relation is vg/c=(dE/dp)/c=1+(α−1)AαEα−2/2+O(Aα2)v_{g}/c=(dE/dp)/c=1+(\alpha-1)A_{\alpha}E^{\alpha-2}/2+O(A_{\alpha}^{2}). The associated length scale is λA≔hc∣Aα∣1/(α−2)\lambda_{A}\coloneqq hc|A_{\alpha}|^{1/(\alpha-2)}, where hh is Planck’s constant. λA\lambda_{A} gives the scale of modifications to the Newtonian potential (the Yukawa potential for α=0\alpha=0) associated with this dispersion relation.

While Eq. (2) is a purely phenomenological model, it encompasses a variety of more fundamental predictions (at least to leading order) Mirshekari et al. 2012; Yunes et al. 2016. In particular, A0>0A_{0}>0 corresponds to a massive graviton, i.e., the same dispersion as for a massive particle in vacuo Will 1998, with a graviton mass given by mg=A01/2/c2m_{g}=A_{0}^{1/2}/c^{2}. Thus, the Yukawa screening length is λ0=h/(mgc)\lambda_{0}=h/(m_{g}c). Furthermore, α\alpha values of 2.52.5, 33, and 44 correspond to the leading predictions of multi-fractal spacetime Calcagni 2010; doubly special relativity Amelino-Camelia 2002; and Hořava-Lifshitz Hořava 2009 and extra dimensional Sefiedgar et al. 2011 theories, respectively. The standard model extension also gives a leading contribution with α=4\alpha=4 Kostelecký and Mewes 2016, only considering the non-birefringent terms; our analysis does not allow for birefringence.

In order to obtain a waveform model with which to constrain these propagation effects, we start by assuming that the waveform extracted in the binary’s local wave zone (i.e., near to the binary compared to the distance from the binary to Earth, but far from the binary compared to its own size) is well-described by a waveform in GR. This is likely to be a good assumption for α<2\alpha<2, where we constrain λA\lambda_{A} to be much larger than the size of the binary. For α>2\alpha>2, where we constrain λA\lambda_{A} to be much smaller than the size of the binary, one has to posit a screening mechanism in order to be able to assume that the waveform in the binary’s local wave zone is well-described by GR, as well as for this model to evade Solar System constraints. Since we are able to bound these propagation effects to be very small, we can work to linear order in AαA_{\alpha} when computing the effects of this dispersion on the frequency-domain GW phasing, The dimensionless parameter controlling the size of the linear correction is Aαfα−2A_{\alpha}f^{\alpha-2}, which is ≲10−18\lesssim 10^{-18} at the 90% credible level for the events we consider and frequencies up to 11 kHz. thus obtaining a correction Mirshekari et al. 2012 that is added to Φ(f)\Phi(f) in Eq. (1):

Here, DLD_{\text{L}} is the binary’s luminosity distance, Mdet\mathcal{M}^{\text{det}} is the binary’s detector-frame (i.e., redshifted) chirp mass, and λA,eff\lambda_{A,\text{eff}} is the effective wavelength parameter used in the sampling, defined as

The parameter zz is the binary’s redshift, and DαD_{\alpha} is a distance parameter given by

where H0=67.90 km s−1 Mpc−1H_{0}=67.90\text{ km s}^{-1}\text{ Mpc}^{-1} is the Hubble constant, and Ωm=0.3065\Omega_{\text{m}}=0.3065 and ΩΛ=0.6935\Omega_{\Lambda}=0.6935 are the matter and dark energy density parameters; these are the TT+lowP+lensing+ext values from Ade et al. 2016. We use these values for consistency with the results presented in Abbott et al. 2019. If we instead use the more recent results from Aghanim et al. 2018, specifically the TT,TE,EE+lowE+lensing+BAO values used for comparison in Abbott et al. 2019, then there are very minor changes to the results presented in this section. For instance, the upper bounds in Table 4 change by at most ∼0.1%\sim 0.1\%.

The dephasing in Eq. (3) is obtained by treating the gravitational wave as a stream of particles (gravitons), which travel at the particle velocity vp/c=pc/E=1−AαEα−2/2+O(Aα2)v_{p}/c=pc/E=1-A_{\alpha}E^{\alpha-2}/2+O(A_{\alpha}^{2}). There are suggestions to use the particle velocity when considering doubly special relativity, though there are also suggestions to use the group velocity vgv_{g} in that case (see, e.g., Amelino-Camelia et al. 2006 and references therein for both arguments). However, the group velocity is appropriate for, e.g., multi-fractal spacetime theories (see, e.g., Calcagni 2017). To convert the bounds presented here to the case where the particles travel at the group velocity, scale the AαA_{\alpha} bounds for α≠1\alpha\neq 1 by factors of 1/(1−α)1/(1-\alpha). The group velocity calculation gives an unobservable constant phase shift for α=1\alpha=1.

We consider the cases of positive and negative AαA_{\alpha} separately, and obtain the results shown in Table 4 and Fig. 5 when applying this analysis to the GW events under consideration. While we sample with a flat prior in log⁡λA,eff\log\lambda_{A,\text{eff}}, our bounds are given using priors flat in AαA_{\alpha} for all results except for the mass of the graviton, where we use a prior flat in the graviton mass. We also show the results from combining together all the signals that satisfy our selection criterion. We are able to combine together the results from different signals with no ambiguity, since the known distance dependence is accounted for in the waveforms.

Figure 6 displays the full AαA_{\alpha} posteriors obtained by combining all selected events (using IMRPhenomPv2 waveforms). To obtain the full AαA_{\alpha} posteriors, we combine together the positive and negative AαA_{\alpha} results for individual events by weighting by their Bayesian evidences; we then combine the posteriors from individual events. We give the analogous plots for the individual events in Sec. A.4 of the Appendix. The combined positive and negative AαA_{\alpha} posteriors are also used to compute the GR quantiles given in Table 4, which give the probability to have Aα<0A_{\alpha}<0, where Aα=0A_{\alpha}=0 is the GR value. Thus, large or small values of the GR quantile indicate that the distribution is not peaked close to the GR value. For a GR signal, the GR quantile will be distributed uniformly in $fordifferentnoiserealizations.TheGRquantileswefindareconsistentwithsuchauniformdistribution.Inparticular,the(two−tailed)metafor different noise realizations. The GR quantiles we find are consistent with such a uniform distribution. In particular, the (two-tailed) metap−valueforalleventsand-value for all events and\alphavaluesobtainedusingFisher’smethodFisher1948(asinSec.V.1)isvalues obtained using Fisher’s method Fisher 1948 (as in Sec. V.1) is0.9995$.

We find that the combined bounds overall improve on those quoted in Abbott et al. 2017 by roughly the factor of 7/3≃1.5\sqrt{7/3}\simeq 1.5 expected from including more events, with the bounds for some quantities improving by up to a factor of 2.52.5, due to the inclusion of several more massive and distant systems in the sample. While the results in Abbott et al. 2017 were affected by a slight normalization issue, and also had insufficiently fine binning in the computation of the upper bounds, we find improvements of up to a factor of 3.43.4 when comparing to the combined GW150914 + GW151226 + GW170104 bounds we compute here. These massive and distant systems, notably GW170823 (and GW170729, which is not included in the combined results), generally give the best individual bounds, particularly for small values of α\alpha, where the dephasing is largest at lower frequencies. Closer and less massive systems such as GW151226 and GW170608 provide weaker bounds, overall. However, their bounds can be comparable to those of the more massive, distant events for larger values of α\alpha. The lighter systems have more power at higher frequencies where the dephasing from the modified dispersion is larger for larger values of α\alpha.

The new combined bound on the mass of the graviton of mg≤4.7×10−23 eV/c2m_{g}\leq 4.7\times 10^{-23}\text{ eV}/c^{2} is a factor of 1.61.6 improvement on the one presented in Abbott et al. 2017. It is also a small improvement on the bound of mg≤6.76×10−23 eV/c2m_{g}\leq 6.76\times 10^{-23}\text{ eV}/c^{2} (90%90\% confidence level) obtained from Solar System ephemerides in Bernus et al. 2019. The much stronger bound in Will 2018 is deduced from a post-fit analysis (i.e., using the residuals of a fit to Solar System ephemerides performed without including the effects of a massive graviton). It may therefore overestimate Solar System constraints, as is indeed seen to be the case in Bernus et al. 2019. However, these bounds are complementary, since the GW bound comes from the radiative sector, while the Solar System bound considers the static modification to the Newtonian potential. See, e.g., de Rham et al. 2017 for a review of bounds on the mass of the graviton.

We find that the posterior on AαA_{\alpha} peaks away from 00 in some cases (illustrated in Sec. A.4 of the Appendix), and the GR quantile is in one of the tails of the distribution. This feature is expected for a few out of 10 events, simply from Gaussian noise fluctuations. We have performed simulations of 100 GR sources with source-frame component masses lying between 25 and 45 M⊙M_{\odot}, isotropically distributed spins with dimensionless magnitudes up to 0.990.99, and at luminosity distances between 500 and 800 Mpc. These simulations used the waveform model IMRPhenomPv2 and considered the Advanced LIGO and Virgo network, using Gaussian noise with the detectors’ design sensitivity power spectral densities. We found that in about 20 – 30% of cases, the GR quantile lies in the tails of the distribution (i.e., <10%<10\% or >90%>90\%), when the sources injected are analyzed using the same waveform model (IMRPhenomPv2).

In order to assess the impact of waveform systematics, we also analyze some events using the aligned-spin SEOBNRv4 model. We consider GW170729 and GW170814 in depth in this study because the GR quantiles of the IMRPhenomPv2 results lie in the tails of the distributions, and find that the 90%90\% upper bounds and GR quantiles presented in Table 4 differ by at most a factor of 2.32.3 for GW170729 and 1.51.5 for GW170814 when computed using the SEOBNRv4 model. These results are presented in Sec. A.4 of the Appendix.

There are also uncertainties in the determination of the 90% bounds due to the finite number of samples and the long tails of the distributions. As in Ref. Abbott et al. 2017, we quantify this uncertainty using Bayesian bootstrapping Rubin 1981. We use 1000 bootstrap realizations for each value of α\alpha and sign of AαA_{\alpha}, obtaining a distribution of 90% bounds on AαA_{\alpha}. We consider the 90% credible interval of this distribution and find that its width is <30%<30\% of the values for the 90% bounds on AαA_{\alpha} given in Table 4 for all but 1010 of the 160160 cases we consider (counting positive and negative AαA_{\alpha} cases separately). For GW170608, A4<0A_{4}<0, the width of the 90% credible interval from bootstrapping is 91%91\% of the value in Table 4. This ratio is ≤47%\leq 47\% for all the remaining cases. Thus, there are a few cases where the bootstrapping uncertainty in the bound on AαA_{\alpha} is large, but for most cases, this is not a substantial uncertainty.

VIII Polarizations

Generic metric theories of gravity may allow up to six polarizations of gravitational waves Eardley et al. 1973: two tensor modes (helicity ±2\pm 2), two vector modes (helicity ±1\pm 1), and two scalar modes (helicity 0). Of these, only the two tensor modes (++ and ×\times) are permitted in GR. We may attempt to reconstruct the polarization content of a passing GW using a network of detectors Will 2014; Chatziioannou et al. 2012; Isi et al. 2017; Callister et al. 2017; Isi and Weinstein 2017. This is possible because instruments with different orientations will respond differently to signals from a given sky location depending on their polarization. In particular, the strain signal in detector II can be written as hI(t)=∑AFAIhA(t)h_{I}(t)=\sum_{A}F^{A}{}_{I}h_{A}(t), with FAIF^{A}{}_{I} the detector’s response function and hA(t)h_{A}(t) the AA-polarized part of the signal Will 2014; Błaut 2012.

In order to fully disentangle the polarization content of a transient signal, at least 5 detectors are needed to break all degeneracies Chatziioannou et al. 2012. Differential-arm detectors are only sensitive to the traceless scalar mode, meaning we can only hope to distinguish five, not six, polarizations. This limits the polarization measurements that are currently feasible. In spite of this, we may extract some polarization information from signals detected with both LIGO detectors and Virgo Isi and Weinstein 2017. This was done previously with GW170814 and GW170817 to provide evidence that GWs are tensor polarized, instead of fully vector or fully scalar Abbott et al. 2017; Abbott et al. 2019. Besides GW170814, there are three binary black hole events that were detected with the full network (GW170729, GW170809, and GW170818). Of these events, only GW170818 has enough SNR and is sufficiently well localized to provide any relevant information (cf. Fig. 8 in Abbott et al. 2019). The Bayes factors (marginalized likelihood ratios) obtained in this case are 12±312\pm 3 for tensor vs vector and 407±100407\pm 100 for tensor vs scalar, where the error corresponds to the uncertainty due to discrete sampling in the evidence computations. These values are comparable to those from GW170814, for which the latest recalibrated and cleaned data (cf. Sec. II) yield Bayes factors of 30±430\pm 4 and 220±27220\pm 27 for tensor vs vector and scalar respectively. These values are less stringent than those previously published in Abbott et al. 2017. This is solely due to the change in data, which impacted the sky locations inferred under the non-GR hypotheses. Values from these binary black holes are many orders of magnitude weaker than those obtained from GW170817, where we benefited from the precise sky-localization provided by an electromagnetic counterpart Abbott et al. 2019.

IX Conclusions and outlook

We have presented the results from various tests of GR performed using the binary black hole signals from the catalog GWTC-1 LIGO Scientific Collaboration and Virgo Collaboration 2018, i.e., those observed by Advanced LIGO and Advanced Virgo during the first two observing runs of the advanced detector era. These tests, which are among the first tests of GR in the highly relativistic, nonlinear regime of strong gravity, do not reveal any inconsistency of our data with the predictions of GR. We have presented full results on four tests of the consistency of the data with gravitational waveforms from binary black hole systems as predicted by GR. The first two of these tests check the self-consistency of our analysis. One checks that the residual remaining after subtracting the best-fit waveform is consistent with detector noise. The other checks that the final mass and spin inferred from the low- and high-frequency parts of the signal are consistent. The third and fourth tests introduce parameterized deviations in the waveform model and check that these deviations are consistent with their GR value of zero. In one test, these deviations are completely phenomenological modifications of the coefficients in a waveform model, including the post-Newtonian coefficients. In the other test, the deviations are those arising from the propagation of GWs with a modified dispersion relation, which includes the dispersion due to a massive graviton as a special case. In addition, we also check whether the observed polarizations are consistent with being purely tensor modes (as expected in GR) as opposed to purely scalar modes or vector modes.

With the expected observations of additional binary black hole merger events in the upcoming LIGO/Virgo observing runs Abbott et al. 2018; Abbott et al. 2019, the statistical errors of the combined results will soon decrease significantly. A number of potential sources of systematic errors (due to imperfect modeling of GR waveforms, calibration uncertainties, noise artifacts, etc.) need to be understood for future high-precision tests of strong gravity using GW observations. However, work to improve the analysis on all these fronts is well underway, for instance the inclusion of full spin-precession dynamics Taracchini et al. 2014; Babak et al. 2017; Blackman et al. 2017; Khan et al. 2019; Varma et al. 2019, non-quadrupolar modes Blackman et al. 2017; London et al. 2018; Cotesta et al. 2018; Varma et al. 2019; Varma et al. 2019, and eccentricity Hinderer and Babak 2017; Cao and Han 2017; Hinder et al. 2018; Huerta et al. 2018; Klein et al. 2018; Moore et al. 2018; Moore and Yunes 2019; Tiwari et al. 2019 in waveform models, as well as analyses that compare directly with numerical relativity waveforms Abbott et al. 2016; Lange et al. 2017. Additionally, a new, more flexible parameter estimation infrastructure is currently being developed Ashton et al. 2019, and this will allow for improvements in, e.g., the treatment of calibration uncertainties or PSD estimation to be incorporated more easily. We thus expect that tests of general relativity using the data from upcoming observing runs will be able to take full advantage of the increased sensitivity of the detectors.

Appendix A Individual results and systematics studies

In the main body of the paper, for most analyses, we present only the combined results from all events. Here we present the posteriors from various tests obtained from individual events. In addition, we offer a limited discussion on systematic errors in the analysis, due to the specific choice of a GR waveform approximant.

As mentioned in Sec. V.1, the residuals test is sensitive to all kinds of disagreement between the best-fit GR-based waveform and the data. This is true whether the disagreement is due to actual deviations from GR or more mundane reasons, like physics missing from our waveform models (e.g., higher-order modes). Had we found compelling evidence of coherent power in the residuals that could not be explained by instrumental noise, further investigations would be required to determine its origin. However, given our null result, we can simply state that we find no evidence for shortcomings in the best-fit waveform, neither from deviations from GR nor modeling systematics.

As the sensitivity of the detectors improves, the issue of systematics will become increasingly more important. To address this, future versions of this test will be carried out by subtracting a best-fit waveform produced with more accurate GR-based models, including numerical relativity.

A.2 Inspiral-merger-ringdown consistency test

A.3 Parametrized tests of gravitational wave generation

Figures 8 and 9 report the parameterized tests of waveform deviations for the individual events, augmenting the results shown in Fig. 3. A statistical summary of the posterior PDFs, showing median and symmetric 90% credible level bounds for the measured parameters is given in Table 5. Sources with low SNR in the inspiral regime yield uninformative posterior distributions on δφ^i\delta\hat{\varphi}_{i}. These sources are the ones further away and with higher mass, which merge at lower frequencies. For instance, although GW170823 has a total mass close to that of GW150914, being much further away (and redshifted to lower frequencies) makes it a low-SNR event, leaving very little information content in the inspiral regime. The same holds true for GW170729, which has a larger mass. Conversely, low-mass events like GW170608, having a significantly larger SNR in the inspiral regime and many more cycles in the frequency band provide very strong constraints in the δφ^i\delta\hat{\varphi}_{i} parameters (especially the low-order ones) while providing no useful constraints in the merger-ringdown parameters δα^i\delta\hat{\alpha}_{i}.

Both here and in Sec. VI we report results on the parametrized deviations in the PN regime using two waveform models, IMRPhenomPv2 and SEOBNRv4. There is a subtle difference between the ways deviations from GR are introduced and parametrized in the two models. With IMRPhenomPv2, we directly constrain δφ^i\delta\hat{\varphi}_{i}, which represent fractional deviations in the non-spinning portion of the (i/2)(i/2)PN phase coefficients. The SEOBNRv4 analysis instead uses a parameterization that also applies the fractional deviations to spin contributions, as described in Abbott et al. 2019. The results are then mapped post-hoc from this native parameterization to posteriors on δφ^i\delta\hat{\varphi}_{i}, shown in Figs. 3 and 8 (black solid lines).

In the SEOBNRv4 analysis at 3.53.5PN, the native (spin-inclusive) posteriors contain tails that extend to the edge of the prior range. This is due to a zero-crossing of the 3.53.5PN term in the (η,a1,a2)(\eta,a_{1},a_{2}) parameter space, which makes the corresponding relative deviation ill-defined. After the post-hoc mapping to posteriors on δφ^7\delta\hat{\varphi}_{7}, no tails appear and we find good agreement with the IMRPhenomPv2 analysis, as expected. By varying the prior range, we estimate a systematic uncertainty of at most a few percent on the quoted 90%90\% bounds due to the truncation of tails.

A.4 Parameterized tests of gravitational wave propagation

Posteriors on AαA_{\alpha} for individual events are shown in Fig. 10, with data for positive and negative AαA_{\alpha} combined into one violin plot. We provide results for all events with the IMRPhenomPv2 waveform model and also show results of the analysis with the SEOBNRv4 waveform model for GW170729 and GW170814. In Table 6 we compare the 90% bounds on AαA_{\alpha} and GR quantiles obtained with IMRPhenomPv2 and SEOBNRv4 for GW170729 and GW170814. We focus on these two events because the GR quantiles obtained with IMRPhenomPv2 lie in the tails of the distributions, and we find that this remains true for most α\alpha values in the analysis with SEOBNRv4. For GW170729 and α∈{0,0.5,1}\alpha\in\{0,0.5,1\}, the GR quantiles obtained using the two waveforms differ by factors of ∼2\sim 2; the two waveforms give values that are in much closer agreement for the other cases.

Additionally, for the GW151012 event and certain α\alpha values, a technical issue with our computation of the likelihood meant that specific points with relatively large values of AαA_{\alpha} had to be manually removed from the posterior distribution. In particular, for computational efficiency, the likelihood is calculated on as short a segment of data as practical, with duration set by the longest waveform to be sampled. Large values of AαA_{\alpha} yield highly dispersed waveforms that are pushed beyond the confines of the segment we use, causing the waveform templates to wrap around the boundaries. This invalidates the assumptions underlying our likelihood computation and causes an artificial enhancement of the SNR as reported by the analysis. As expected, recomputing the SNRs for these points on a segment that properly fits the waveform results in smaller values that are consistent with noise. Therefore, we exclude from our analysis parameter values yielding waveforms that would not be contained by the data segment used, which is equivalent to using a stricter prior on AαA_{\alpha}. Failure to do this may result in the appearance of outliers with spuriously high likelihood for large values of AαA_{\alpha}, as we have seen in our own analysis.

References