p7cs.NE

Category

cs.NE

741 papers

Show Your Work: Scratchpads for Intermediate Computation with Language Models

Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, Augustus Odena

5 annotations

2112.00114

Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, Jeff Dean

1701.06538

A Structured Self-attentive Sentence Embedding

Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, Yoshua Bengio

1703.03130

Quasi-Recurrent Neural Networks

James Bradbury, Stephen Merity, Caiming Xiong, Richard Socher

1611.01576

Sequential Short-Text Classification with Recurrent and Convolutional Neural Networks

Ji Young Lee, Franck Dernoncourt

1603.03827

Long Short-Term Memory-Networks for Machine Reading

Jianpeng Cheng, Li Dong, Mirella Lapata

1601.06733

Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations

David Krueger, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Aaron Courville, Chris Pal

1606.01305

Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks

Tim Salimans, Diederik P. Kingma

1602.07868

Recurrent Highway Networks

Julian Georg Zilly, Rupesh Kumar Srivastava, Jan Koutník, Jürgen Schmidhuber

1607.03474

Pixel Recurrent Neural Networks

Aaron van den Oord, Nal Kalchbrenner, Koray Kavukcuoglu

1601.06759

Listen, Attend and Spell

William Chan, Navdeep Jaitly, Quoc V. Le, Oriol Vinyals

1508.01211

Neural Architecture Search with Reinforcement Learning

Barret Zoph, Quoc V. Le

1611.01578

Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, Yoshua Bengio

1406.1078

Neural Machine Translation by Jointly Learning to Align and Translate

Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio

1409.0473

Predicting Deep Zero-Shot Convolutional Neural Networks using Textual Descriptions

Jimmy Ba, Kevin Swersky, Sanja Fidler, Ruslan Salakhutdinov

1506.00511

Memory-Efficient Backpropagation Through Time

Audrūnas Gruslys, Remi Munos, Ivo Danihelka, Marc Lanctot, Alex Graves

1606.03401

Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, Jianfeng Gao

2203.03466

Regularizing and Optimizing LSTM Language Models

Stephen Merity, Nitish Shirish Keskar, Richard Socher

1708.02182

Tensor Programs IIb: Architectural Universality of Neural Tangent Kernel Training Dynamics

Greg Yang, Etai Littwin

2105.03703

Dynamic Evaluation of Neural Sequence Models

Ben Krause, Emmanuel Kahembwe, Iain Murray, Steve Renals

1709.07432

Transformer Quality in Linear Time

Weizhe Hua, Zihang Dai, Hanxiao Liu, Quoc V. Le

2202.10447

Character-Aware Neural Language Models

Yoon Kim, Yacine Jernite, David Sontag, Alexander M. Rush

1508.06615

GLU Variants Improve Transformer

Noam Shazeer

2002.05202

Highway and Residual Networks learn Unrolled Iterative Estimation

Klaus Greff, Rupesh K. Srivastava, Jürgen Schmidhuber

1612.07771

End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results

Jan Chorowski, Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio

1412.1602

Simple Recurrent Units for Highly Parallelizable Recurrence

Tao Lei, Yu Zhang, Sida I. Wang, Hui Dai, Yoav Artzi

1709.02755

DRAW: A Recurrent Neural Network For Image Generation

Karol Gregor, Ivo Danihelka, Alex Graves, Danilo Jimenez Rezende, Daan Wierstra

1502.04623

Pointing the Unknown Words

Caglar Gulcehre, Sungjin Ahn, Ramesh Nallapati, Bowen Zhou, Yoshua Bengio

1603.08148

Noisy Activation Functions

Caglar Gulcehre, Marcin Moczulski, Misha Denil, Yoshua Bengio

1603.00391

Reasoning about Entailment with Neural Attention

Tim Rocktäschel, Edward Grefenstette, Karl Moritz Hermann, Tomáš Kočiský, Phil Blunsom

1509.06664

Fast and Accurate Recurrent Neural Network Acoustic Models for Speech Recognition

Haşim Sak, Andrew Senior, Kanishka Rao, Françoise Beaufays

1507.06947

Training Very Deep Networks

Rupesh Kumar Srivastava, Klaus Greff, Jürgen Schmidhuber

1507.06228

Teaching Machines to Read and Comprehend

Karl Moritz Hermann, Tomáš Kočiský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, Phil Blunsom

1506.03340

Building End-To-End Dialogue Systems Using Generative Hierarchical Neural Network Models

Iulian V. Serban, Alessandro Sordoni, Yoshua Bengio, Aaron Courville, Joelle Pineau

1507.04808

Learning to See by Moving

Pulkit Agrawal, Joao Carreira, Jitendra Malik

1505.01596

Investigating gated recurrent neural networks for speech synthesis

Zhizheng Wu, Simon King

1601.02539

Attention-Based Models for Speech Recognition

Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, Yoshua Bengio

1506.07503

Speech Recognition with Deep Recurrent Neural Networks

Alex Graves, Abdel-rahman Mohamed, Geoffrey Hinton

1303.5778

Highway Networks

Rupesh Kumar Srivastava, Klaus Greff, Jürgen Schmidhuber

1505.00387

Incorporating Copying Mechanism in Sequence-to-Sequence Learning

Jiatao Gu, Zhengdong Lu, Hang Li, Victor O. K. Li

1603.06393

Pointer Networks

Oriol Vinyals, Meire Fortunato, Navdeep Jaitly

1506.03134

Learning Longer Memory in Recurrent Neural Networks

Tomas Mikolov, Armand Joulin, Sumit Chopra, Michael Mathieu, Marc'Aurelio Ranzato

1412.7753

Improving neural networks by preventing co-adaptation of feature detectors

Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, Ruslan R. Salakhutdinov

1207.0580

Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs

Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, Alan L. Yuille

1412.7062

Distilling the Knowledge in a Neural Network

Geoffrey Hinton, Oriol Vinyals, Jeff Dean

1503.02531

FitNets: Hints for Thin Deep Nets

Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, Yoshua Bengio

1412.6550

Compressing Neural Networks with the Hashing Trick

Wenlin Chen, James T. Wilson, Stephen Tyree, Kilian Q. Weinberger, Yixin Chen

1504.04788

LSTM: A Search Space Odyssey

Klaus Greff, Rupesh Kumar Srivastava, Jan Koutník, Bas R. Steunebrink, Jürgen Schmidhuber

1503.04069

Neural Responding Machine for Short-Text Conversation

Lifeng Shang, Zhengdong Lu, Hang Li

1503.02364

Unsupervised Learning of Video Representations using LSTMs

Nitish Srivastava, Elman Mansimov, Ruslan Salakhutdinov

1502.04681

Neural Arithmetic Logic Units

Andrew Trask, Felix Hill, Scott Reed, Jack Rae, Chris Dyer, Phil Blunsom

1808.00508

Recurrent Neural Network Regularization

Wojciech Zaremba, Ilya Sutskever, Oriol Vinyals

1409.2329

Caffe: Convolutional Architecture for Fast Feature Embedding

Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, Trevor Darrell

1408.5093

Neural Turing Machines

Alex Graves, Greg Wayne, Ivo Danihelka

1410.5401

Deep Double Descent: Where Bigger Models and More Data Hurt

Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, Ilya Sutskever

1912.02292

A Neural Network Approach to Context-Sensitive Generation of Conversational Responses

Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, Bill Dolan

1506.06714

Binary Neural Networks: A Survey

Haotong Qin, Ruihao Gong, Xianglong Liu, Xiao Bai, Jingkuan Song, Nicu Sebe

2004.03333

Toxicity Prediction using Deep Learning

Thomas Unterthiner, Andreas Mayr, Günter Klambauer, Sepp Hochreiter

1503.01445

Learning to Compose Neural Networks for Question Answering

Jacob Andreas, Marcus Rohrbach, Trevor Darrell, Dan Klein

1601.01705

Parameter Space Noise for Exploration

Matthias Plappert, Rein Houthooft, Prafulla Dhariwal, Szymon Sidor, Richard Y. Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, Marcin Andrychowicz

1706.01905

Deep Biaffine Attention for Neural Dependency Parsing

Timothy Dozat, Christopher D. Manning

1611.01734

Natural Neural Networks

Guillaume Desjardins, Karen Simonyan, Razvan Pascanu, Koray Kavukcuoglu

1507.00210

Learning Continuous Control Policies by Stochastic Value Gradients

Nicolas Heess, Greg Wayne, David Silver, Timothy Lillicrap, Yuval Tassa, Tom Erez

1510.09142

Learning to Answer Questions From Image Using Convolutional Neural Network

Lin Ma, Zhengdong Lu, Hang Li

1506.00333

How transferable are features in deep neural networks?

Jason Yosinski, Jeff Clune, Yoshua Bengio, Hod Lipson

1411.1792

Neural Module Networks

Jacob Andreas, Marcus Rohrbach, Trevor Darrell, Dan Klein

1511.02799

Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering

Huijuan Xu, Kate Saenko

1511.05234

Regularized Evolution for Image Classifier Architecture Search

Esteban Real, Alok Aggarwal, Yanping Huang, Quoc V Le

1802.01548

Efficient Neural Architecture Search via Parameter Sharing

Hieu Pham, Melody Y. Guan, Barret Zoph, Quoc V. Le, Jeff Dean

1802.03268

Deeply-Supervised Nets

Chen-Yu Lee, Saining Xie, Patrick Gallagher, Zhengyou Zhang, Zhuowen Tu

1409.5185

Addressing the Rare Word Problem in Neural Machine Translation

Minh-Thang Luong, Ilya Sutskever, Quoc V. Le, Oriol Vinyals, Wojciech Zaremba

1410.8206

Massively Multitask Networks for Drug Discovery

Bharath Ramsundar, Steven Kearnes, Patrick Riley, Dale Webster, David Konerding, Vijay Pande

1502.02072

Improved Variational Autoencoders for Text Modeling using Dilated Convolutions

Zichao Yang, Zhiting Hu, Ruslan Salakhutdinov, Taylor Berg-Kirkpatrick

1702.08139

Learning to Execute

Wojciech Zaremba, Ilya Sutskever

1410.4615

Multi-task Neural Networks for QSAR Predictions

George E. Dahl, Navdeep Jaitly, Ruslan Salakhutdinov

1406.1231

Generating Sequences With Recurrent Neural Networks

Alex Graves

1308.0850

Adaptive Neural Networks for Efficient Inference

Tolga Bolukbasi, Joseph Wang, Ofer Dekel, Venkatesh Saligrama

1702.07811

Spectral Networks and Locally Connected Networks on Graphs

Joan Bruna, Wojciech Zaremba, Arthur Szlam, Yann LeCun

1312.6203

Visual7W: Grounded Question Answering in Images

Yuke Zhu, Oliver Groth, Michael Bernstein, Li Fei-Fei

1511.03416

Ask Me Anything: Dynamic Memory Networks for Natural Language Processing

Ankit Kumar, Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury, Ishaan Gulrajani, Victor Zhong, Romain Paulus, Richard Socher

1506.07285

Reinforcement Learning with Unsupervised Auxiliary Tasks

Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, Koray Kavukcuoglu

1611.05397

Dynamic Memory Networks for Visual and Textual Question Answering

Caiming Xiong, Stephen Merity, Richard Socher

1603.01417

Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, Ilya Sutskever

1703.03864

The Power of Depth for Feedforward Neural Networks

Ronen Eldan, Ohad Shamir

1512.03965

A Style-Based Generator Architecture for Generative Adversarial Networks

Tero Karras, Samuli Laine, Timo Aila

1812.04948

Analyzing and Improving the Image Quality of StyleGAN

Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, Timo Aila

1912.04958

A Spectral Approach to Gradient Estimation for Implicit Distributions

Jiaxin Shi, Shengyang Sun, Jun Zhu

1806.02925

A Neural Representation of Sketch Drawings

David Ha, Douglas Eck

1704.03477

Hierarchical Representations for Efficient Architecture Search

Hanxiao Liu, Karen Simonyan, Oriol Vinyals, Chrisantha Fernando, Koray Kavukcuoglu

1711.00436

Sensitivity and Generalization in Neural Networks: an Empirical Study

Roman Novak, Yasaman Bahri, Daniel A. Abolafia, Jeffrey Pennington, Jascha Sohl-Dickstein

1802.08760

The NarrativeQA Reading Comprehension Challenge

Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, Edward Grefenstette

1712.07040

Network In Network

Min Lin, Qiang Chen, Shuicheng Yan

1312.4400

Size-Independent Sample Complexity of Neural Networks

Noah Golowich, Alexander Rakhlin, Ohad Shamir

1712.06541

Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks

Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, Ruosong Wang

1901.08584

The Evolved Transformer

David R. So, Chen Liang, Quoc V. Le

1901.11117

On the importance of single directions for generalization

Ari S. Morcos, David G. T. Barrett, Neil C. Rabinowitz, Matthew Botvinick

1803.06959

Learning and Generalization in Overparameterized Neural Networks, Going Beyond Two Layers

Zeyuan Allen-Zhu, Yuanzhi Li, Yingyu Liang

1811.04918

Deep Learning with Dynamic Computation Graphs

Moshe Looks, Marcello Herreshoff, DeLesley Hutchins, Peter Norvig

1702.02181

Learning values across many orders of magnitude

Hado van Hasselt, Arthur Guez, Matteo Hessel, Volodymyr Mnih, David Silver

1602.07714

Learning to learn by gradient descent by gradient descent

Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W. Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, Nando de Freitas

1606.04474

RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning

Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, Pieter Abbeel

1611.02779

Towards Accurate Generative Models of Video: A New Metric & Challenges

Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, Sylvain Gelly

1812.01717

Learning to Generate Reviews and Discovering Sentiment

Alec Radford, Rafal Jozefowicz, Ilya Sutskever

1704.01444

Primer: Searching for Efficient Transformers for Language Modeling

David R. So, Wojciech Mańke, Hanxiao Liu, Zihang Dai, Noam Shazeer, Quoc V. Le

2109.08668

Convolutional Neural Network Architectures for Matching Natural Language Sentences

Baotian Hu, Zhengdong Lu, Hang Li, Qingcai Chen

1503.03244

Liquid Structural State-Space Models

Ramin Hasani, Mathias Lechner, Tsun-Hsuan Wang, Makram Chahine, Alexander Amini, Daniela Rus

2209.12951

Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit

Boaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham Kakade, Eran Malach, Cyril Zhang

2207.08799

Iterative Alternating Neural Attention for Machine Reading

Alessandro Sordoni, Philip Bachman, Adam Trischler, Yoshua Bengio

1606.02245

Symbolic Discovery of Optimization Algorithms

Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Yao Liu, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, Quoc V. Le

2302.06675

Predicting Parameters in Deep Learning

Misha Denil, Babak Shakibi, Laurent Dinh, Marc'Aurelio Ranzato, Nando de Freitas

1306.0543

Memory Augmented Neural Networks with Wormhole Connections

Caglar Gulcehre, Sarath Chandar, Yoshua Bengio

1701.08718

Stochastic Pooling for Regularization of Deep Convolutional Neural Networks

Matthew D. Zeiler, Rob Fergus

1301.3557

Spectrally-normalized margin bounds for neural networks

Peter Bartlett, Dylan J. Foster, Matus Telgarsky

1706.08498

Transition-Based Dependency Parsing with Stack Long Short-Term Memory

Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, Noah A. Smith

1505.08075

Bridging the Gaps Between Residual Learning, Recurrent Neural Networks and Visual Cortex

Qianli Liao, Tomaso Poggio

1604.03640

Exploring Hidden Dimensions in Parallelizing Convolutional Neural Networks

Zhihao Jia, Sina Lin, Charles R. Qi, Alex Aiken

1802.04924

Discrete Event, Continuous Time RNNs

Michael C. Mozer, Denis Kazakov, Robert V. Lindsey

1710.04110

Wide Residual Networks

Sergey Zagoruyko, Nikos Komodakis

1605.07146

Large-Scale Evolution of Image Classifiers

Esteban Real, Sherry Moore, Andrew Selle, Saurabh Saxena, Yutaka Leon Suematsu, Jie Tan, Quoc Le, Alex Kurakin

1703.01041

Towards a Human-like Open-Domain Chatbot

Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, Quoc V. Le

2001.09977

Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights

Aojun Zhou, Anbang Yao, Yiwen Guo, Lin Xu, Yurong Chen

1702.03044

Decoupled Weight Decay Regularization

Ilya Loshchilov, Frank Hutter

1711.05101

Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation

Greg Yang

1902.04760

A Mean Field Theory of Batch Normalization

Greg Yang, Jeffrey Pennington, Vinay Rao, Jascha Sohl-Dickstein, Samuel S. Schoenholz

1902.08129

Deep Big Simple Neural Nets Excel on Handwritten Digit Recognition

Dan Claudiu Ciresan, Ueli Meier, Luca Maria Gambardella, Juergen Schmidhuber

1003.0358

Blocks and Fuel: Frameworks for deep learning

Bart van Merriënboer, Dzmitry Bahdanau, Vincent Dumoulin, Dmitriy Serdyuk, David Warde-Farley, Jan Chorowski, Yoshua Bengio

1506.00619

Fully Convolutional Multi-Class Multiple Instance Learning

Deepak Pathak, Evan Shelhamer, Jonathan Long, Trevor Darrell

1412.7144

Globally Normalized Transition-Based Neural Networks

Daniel Andor, Chris Alberti, David Weiss, Aliaksei Severyn, Alessandro Presta, Kuzman Ganchev, Slav Petrov, Michael Collins

1603.06042

Progressive Growing of GANs for Improved Quality, Stability, and Variation

Tero Karras, Timo Aila, Samuli Laine, Jaakko Lehtinen

1710.10196

Norm-Based Capacity Control in Neural Networks

Behnam Neyshabur, Ryota Tomioka, Nathan Srebro

1503.00036

SGDR: Stochastic Gradient Descent with Warm Restarts

Ilya Loshchilov, Frank Hutter

1608.03983

A Hierarchical Latent Variable Encoder-Decoder Model for Generating Dialogues

Iulian Vlad Serban, Alessandro Sordoni, Ryan Lowe, Laurent Charlin, Joelle Pineau, Aaron Courville, Yoshua Bengio

1605.06069

How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation

Chia-Wei Liu, Ryan Lowe, Iulian V. Serban, Michael Noseworthy, Laurent Charlin, Joelle Pineau

1603.08023

Generalization in Deep Learning

Kenji Kawaguchi, Leslie Pack Kaelbling, Yoshua Bengio

1710.05468

Deep Convolutional Networks on Graph-Structured Data

Mikael Henaff, Joan Bruna, Yann LeCun

1506.05163

Gated Graph Sequence Neural Networks

Yujia Li, Daniel Tarlow, Marc Brockschmidt, Richard Zemel

1511.05493

Improved Techniques for Training GANs

Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, Xi Chen

1606.03498

Human Pose Estimation with Iterative Error Feedback

Joao Carreira, Pulkit Agrawal, Katerina Fragkiadaki, Jitendra Malik

1507.06550

A Neural Algorithm of Artistic Style

Leon A. Gatys, Alexander S. Ecker, Matthias Bethge

1508.06576

Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes

Roman Novak, Lechao Xiao, Jaehoon Lee, Yasaman Bahri, Greg Yang, Jiri Hron, Daniel A. Abolafia, Jeffrey Pennington, Jascha Sohl-Dickstein

1810.05148

Learning from Simulated and Unsupervised Images through Adversarial Training

Ashish Shrivastava, Tomas Pfister, Oncel Tuzel, Josh Susskind, Wenda Wang, Russ Webb

1612.07828

On the Convergence Rate of Training Recurrent Neural Networks

Zeyuan Allen-Zhu, Yuanzhi Li, Zhao Song

1810.12065

Can SGD Learn Recurrent Neural Networks with Provable Generalization?

Zeyuan Allen-Zhu, Yuanzhi Li

1902.01028

Hyperbolic Attention Networks

Caglar Gulcehre, Misha Denil, Mateusz Malinowski, Ali Razavi, Razvan Pascanu, Karl Moritz Hermann, Peter Battaglia, Victor Bapst, David Raposo, Adam Santoro, Nando de Freitas

1805.09786

Neural CRF Parsing

Greg Durrett, Dan Klein

1507.03641

Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation

Albert Gatt, Emiel Krahmer

1703.09902

Learned Optimizers that Scale and Generalize

Olga Wichrowska, Niru Maheswaranathan, Matthew W. Hoffman, Sergio Gomez Colmenarejo, Misha Denil, Nando de Freitas, Jascha Sohl-Dickstein

1703.04813

SGD on Neural Networks Learns Functions of Increasing Complexity

Preetum Nakkiran, Gal Kaplun, Dimitris Kalimeris, Tristan Yang, Benjamin L. Edelman, Fred Zhang, Boaz Barak

1905.11604

LSTMVis: A Tool for Visual Analysis of Hidden State Dynamics in Recurrent Neural Networks

Hendrik Strobelt, Sebastian Gehrmann, Hanspeter Pfister, Alexander M. Rush

1606.07461

Rationalizing Neural Predictions

Tao Lei, Regina Barzilay, Tommi Jaakkola

1606.04155

Exploiting Cyclic Symmetry in Convolutional Neural Networks

Sander Dieleman, Jeffrey De Fauw, Koray Kavukcuoglu

1602.02660

SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks

Angshuman Parashar, Minsoo Rhu, Anurag Mukkara, Antonio Puglielli, Rangharajan Venkatesan, Brucek Khailany, Joel Emer, Stephen W. Keckler, William J. Dally

1708.04485

Benefits of depth in neural networks

Matus Telgarsky

1602.04485

Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations

Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, Yoshua Bengio

1609.07061

Error bounds for approximations with deep ReLU networks

Dmitry Yarotsky

1610.01145

On the Expressive Power of Deep Learning: A Tensor Analysis

Nadav Cohen, Or Sharir, Amnon Shashua

1509.05009

Why Deep Neural Networks for Function Approximation?

Shiyu Liang, R. Srikant

1610.04161

Fast Convolutional Nets With fbfft: A GPU Performance Evaluation

Nicolas Vasilache, Jeff Johnson, Michael Mathieu, Soumith Chintala, Serkan Piantino, Yann LeCun

1412.7580

Deep Neural Networks are Easily Fooled: High Confidence Predictions for Unrecognizable Images

Anh Nguyen, Jason Yosinski, Jeff Clune

1412.1897

Deep Networks with Stochastic Depth

Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, Kilian Weinberger

1603.09382

Intriguing properties of neural networks

Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, Rob Fergus

1312.6199

Convolutional Networks on Graphs for Learning Molecular Fingerprints

David Duvenaud, Dougal Maclaurin, Jorge Aguilera-Iparraguirre, Rafael Gómez-Bombarelli, Timothy Hirzel, Alán Aspuru-Guzik, Ryan P. Adams

1509.09292

Identity Matters in Deep Learning

Moritz Hardt, Tengyu Ma

1611.04231

Deep Voice: Real-time Neural Text-to-Speech

Sercan O. Arik, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, Shubho Sengupta, Mohammad Shoeybi

1702.07825

Towards Deep Neural Network Architectures Robust to Adversarial Examples

Shixiang Gu, Luca Rigazio

1412.5068

End-to-End Differentiable Proving

Tim Rocktäschel, Sebastian Riedel

1705.11040

Fast Training of Convolutional Networks through FFTs

Michael Mathieu, Mikael Henaff, Yann LeCun

1312.5851

The Ubuntu Dialogue Corpus: A Large Dataset for Research in Unstructured Multi-Turn Dialogue Systems

Ryan Lowe, Nissan Pow, Iulian Serban, Joelle Pineau

1506.08909

Object detection via a multi-region & semantic segmentation-aware CNN model

Spyros Gidaris, Nikos Komodakis

1505.01749

Image Super-Resolution Using Deep Convolutional Networks

Chao Dong, Chen Change Loy, Kaiming He, Xiaoou Tang

1501.00092

A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay

Leslie N. Smith

1803.09820

Recent Advances in Recurrent Neural Networks

Hojjat Salehinejad, Sharan Sankar, Joseph Barfett, Errol Colak, Shahrokh Valaee

1801.01078

Assessing the Scalability of Biologically-Motivated Deep Learning Algorithms and Architectures

Sergey Bartunov, Adam Santoro, Blake A. Richards, Luke Marris, Geoffrey E. Hinton, Timothy Lillicrap

1807.04587

Rotation-invariant convolutional neural networks for galaxy morphology prediction

Sander Dieleman, Kyle W. Willett, Joni Dambre

1503.07077

Semi-Supervised Learning with Ladder Networks

Antti Rasmus, Harri Valpola, Mikko Honkala, Mathias Berglund, Tapani Raiko

1507.02672

Not Just a Black Box: Learning Important Features Through Propagating Activation Differences

Avanti Shrikumar, Peyton Greenside, Anna Shcherbina, Anshul Kundaje

1605.01713

Adversarial Feature Learning

Jeff Donahue, Philipp Krähenbühl, Trevor Darrell

1605.09782

Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Song Han, Huizi Mao, William J. Dally

1510.00149

SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation

Vijay Badrinarayanan, Alex Kendall, Roberto Cipolla

1511.00561

Learning Activation Functions to Improve Deep Neural Networks

Forest Agostinelli, Matthew Hoffman, Peter Sadowski, Pierre Baldi

1412.6830

On the Number of Linear Regions of Deep Neural Networks

Guido Montúfar, Razvan Pascanu, Kyunghyun Cho, Yoshua Bengio

1402.1869

A Network-based End-to-End Trainable Task-oriented Dialogue System

Tsung-Hsien Wen, David Vandyke, Nikola Mrksic, Milica Gasic, Lina M. Rojas-Barahona, Pei-Hao Su, Stefan Ultes, Steve Young

1604.04562

Deep Networks with Internal Selective Attention through Feedback Connections

Marijn Stollenga, Jonathan Masci, Faustino Gomez, Juergen Schmidhuber

1407.3068

MatConvNet - Convolutional Neural Networks for MATLAB

Andrea Vedaldi, Karel Lenc

1412.4564

Representation Benefits of Deep Feedforward Networks

Matus Telgarsky

1509.08101

Convolutional Rectifier Networks as Generalized Tensor Decompositions

Nadav Cohen, Amnon Shashua

1603.00162

cuDNN: Efficient Primitives for Deep Learning

Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, Evan Shelhamer

1410.0759

Regularization for Deep Learning: A Taxonomy

Jan Kukačka, Vladimir Golkov, Daniel Cremers

1710.10686

Stochastic Variance Reduction for Nonconvex Optimization

Sashank J. Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, Alex Smola

1603.06160

Biologically Motivated Algorithms for Propagating Local Target Representations

Alexander G. Ororbia, Ankur Mali

1805.11703

Stacked What-Where Auto-encoders

Junbo Zhao, Michael Mathieu, Ross Goroshin, Yann LeCun

1506.02351

Domain-Adversarial Training of Neural Networks

Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, Victor Lempitsky

1505.07818

Distribution-Specific Hardness of Learning Neural Networks

Ohad Shamir

1609.01037

Deep Convolutional Inverse Graphics Network

Tejas D. Kulkarni, Will Whitney, Pushmeet Kohli, Joshua B. Tenenbaum

1503.03167

Fast Algorithms for Convolutional Neural Networks

Andrew Lavin, Scott Gray

1509.09308

Learning both Weights and Connections for Efficient Neural Networks

Song Han, Jeff Pool, John Tran, William J. Dally

1506.02626

On the number of response regions of deep feed forward networks with piece-wise linear activations

Razvan Pascanu, Guido Montufar, Yoshua Bengio

1312.6098

Regularizing Neural Networks by Penalizing Confident Output Distributions

Gabriel Pereyra, George Tucker, Jan Chorowski, Łukasz Kaiser, Geoffrey Hinton

1701.06548

Improving Deep Neural Networks with Probabilistic Maxout Units

Jost Tobias Springenberg, Martin Riedmiller

1312.6116

Recurrent Neural Network Grammars

Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, Noah A. Smith

1602.07776

Convolutional Networks for Fast, Energy-Efficient Neuromorphic Computing

Steven K. Esser, Paul A. Merolla, John V. Arthur, Andrew S. Cassidy, Rathinakumar Appuswamy, Alexander Andreopoulos, David J. Berg, Jeffrey L. McKinstry, Timothy Melano, Davis R. Barch, Carmelo di Nolfo, Pallab Datta, Arnon Amir, Brian Taba, Myron D. Flickner, Dharmendra S. Modha

1603.08270

Latent Predictor Networks for Code Generation

Wang Ling, Edward Grefenstette, Karl Moritz Hermann, Tomáš Kočiský, Andrew Senior, Fumin Wang, Phil Blunsom

1603.06744

The Shattered Gradients Problem: If resnets are the answer, then what is the question?

David Balduzzi, Marcus Frean, Lennox Leary, JP Lewis, Kurt Wan-Duo Ma, Brian McWilliams

1702.08591

Learning What and Where to Draw

Scott Reed, Zeynep Akata, Santosh Mohan, Samuel Tenka, Bernt Schiele, Honglak Lee

1610.02454

All You Need is Beyond a Good Init: Exploring Better Solution for Training Extremely Deep Convolutional Neural Networks with Orthonormality and Modulation

Di Xie, Jiang Xiong, Shiliang Pu

1703.01827

Variance Reduction for Faster Non-Convex Optimization

Zeyuan Allen-Zhu, Elad Hazan

1603.05643

Curriculum Dropout

Pietro Morerio, Jacopo Cavazza, Riccardo Volpi, Rene Vidal, Vittorio Murino

1703.06229

Universal Adversarial Perturbations Against Semantic Image Segmentation

Jan Hendrik Metzen, Mummadi Chaithanya Kumar, Thomas Brox, Volker Fischer

1704.05712

Bounding and Counting Linear Regions of Deep Neural Networks

Thiago Serra, Christian Tjandraatmadja, Srikumar Ramalingam

1711.02114

End to End Learning for Self-Driving Cars

Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D. Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, Karol Zieba

1604.07316

Neural GPUs Learn Algorithms

Łukasz Kaiser, Ilya Sutskever

1511.08228

Neural Random-Access Machines

Karol Kurach, Marcin Andrychowicz, Ilya Sutskever

1511.06392

Towards Dropout Training for Convolutional Neural Networks

Haibing Wu, Xiaodong Gu

1512.00242

Neuromorphic Deep Learning Machines

Emre Neftci, Charles Augustine, Somnath Paul, Georgios Detorakis

1612.05596

Unbiased Online Recurrent Optimization

Corentin Tallec, Yann Ollivier

1702.05043

Attention-over-Attention Neural Networks for Reading Comprehension

Yiming Cui, Zhipeng Chen, Si Wei, Shijin Wang, Ting Liu, Guoping Hu

1607.04423

Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks

Chelsea Finn, Pieter Abbeel, Sergey Levine

1703.03400

Variable Rate Image Compression with Recurrent Neural Networks

George Toderici, Sean M. O'Malley, Sung Jin Hwang, Damien Vincent, David Minnen, Shumeet Baluja, Michele Covell, Rahul Sukthankar

1511.06085

Learning to Transduce with Unbounded Memory

Edward Grefenstette, Karl Moritz Hermann, Mustafa Suleyman, Phil Blunsom

1506.02516

Domain Adaptive Neural Networks for Object Recognition

Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang

1409.6041

Inferring Algorithmic Patterns with Stack-Augmented Recurrent Nets

Armand Joulin, Tomas Mikolov

1503.01007

Implicit Regularization in Deep Matrix Factorization

Sanjeev Arora, Nadav Cohen, Wei Hu, Yuping Luo

1905.13655

In-Datacenter Performance Analysis of a Tensor Processing Unit

Norman P. Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara Vazir Ghaemmaghami, Rajendra Gottipati, William Gulland, Robert Hagmann, C. Richard Ho, Doug Hogberg, John Hu, Robert Hundt, Dan Hurt, Julian Ibarz, Aaron Jaffey, Alek Jaworski, Alexander Kaplan, Harshit Khaitan, Andy Koch, Naveen Kumar, Steve Lacy, James Laudon, James Law, Diemthu Le, Chris Leary, Zhuyuan Liu, Kyle Lucke, Alan Lundin, Gordon MacKean, Adriana Maggiore, Maire Mahony, Kieran Miller, Rahul Nagarajan, Ravi Narayanaswami, Ray Ni, Kathy Nix, Thomas Norrie, Mark Omernick, Narayana Penukonda, Andy Phelps, Jonathan Ross, Matt Ross, Amir Salek, Emad Samadiani, Chris Severn, Gregory Sizikov, Matthew Snelham, Jed Souter, Dan Steinberg, Andy Swing, Mercedes Tan, Gregory Thorson, Bo Tian, Horia Toma, Erick Tuttle, Vijay Vasudevan, Richard Walter, Walter Wang, Eric Wilcox, Doe Hyun Yoon

1704.04760

Parallel Multiscale Autoregressive Density Estimation

Scott Reed, Aäron van den Oord, Nal Kalchbrenner, Sergio Gómez Colmenarejo, Ziyu Wang, Dan Belov, Nando de Freitas

1703.03664

Improved Precision and Recall Metric for Assessing Generative Models

Tuomas Kynkäänniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, Timo Aila

1904.06991

ALICE: Towards Understanding Adversarial Learning for Joint Distribution Matching

Chunyuan Li, Hao Liu, Changyou Chen, Yunchen Pu, Liqun Chen, Ricardo Henao, Lawrence Carin

1709.01215

Convolutional Neural Networks for Sentence Classification

Yoon Kim

1408.5882

GraphVAE: Towards Generation of Small Graphs Using Variational Autoencoders

Martin Simonovsky, Nikos Komodakis

1802.03480

Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks

Greg Yang, Dingli Yu, Chen Zhu, Soufiane Hayou

2310.02244

Qualitatively characterizing neural network optimization problems

Ian J. Goodfellow, Oriol Vinyals, Andrew M. Saxe

1412.6544

Neural networks and rational functions

Matus Telgarsky

1706.03301

The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Jonathan Frankle, Michael Carbin

1803.03635

Sparse DNNs with Improved Adversarial Robustness

Yiwen Guo, Chao Zhang, Changshui Zhang, Yurong Chen

1810.09619

Distillation as a Defense to Adversarial Perturbations against Deep Neural Networks

Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, Ananthram Swami

1511.04508

Deep Metric Learning for Practical Person Re-Identification

Dong Yi, Zhen Lei, Stan Z. Li

1407.4979

Learning to Generate Chairs, Tables and Cars with Convolutional Networks

Alexey Dosovitskiy, Jost Tobias Springenberg, Maxim Tatarchenko, Thomas Brox

1411.5928

Generating Images with Perceptual Similarity Metrics based on Deep Networks

Alexey Dosovitskiy, Thomas Brox

1602.02644

The Limitations of Deep Learning in Adversarial Settings

Nicolas Papernot, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z. Berkay Celik, Ananthram Swami

1511.07528

Learning to Compare Image Patches via Convolutional Neural Networks

Sergey Zagoruyko, Nikos Komodakis

1504.03641

Stereo Matching by Training a Convolutional Neural Network to Compare Image Patches

Jure Žbontar, Yann LeCun

1510.05970

Training for Faster Adversarial Robustness Verification via Inducing ReLU Stability

Kai Y. Xiao, Vincent Tjeng, Nur Muhammad Shafiullah, Aleksander Madry

1809.03008

On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport

Lenaic Chizat, Francis Bach

1805.09545

Simple, Efficient, and Neural Algorithms for Sparse Coding

Sanjeev Arora, Rong Ge, Tengyu Ma, Ankur Moitra

1503.00778

Neural Tangent Kernel: Convergence and Generalization in Neural Networks

Arthur Jacot, Franck Gabriel, Clément Hongler

1806.07572

Computing the Stereo Matching Cost with a Convolutional Neural Network

Jure Žbontar, Yann LeCun

1409.4326

Robustness May Be at Odds with Accuracy

Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, Aleksander Madry

1805.12152

Hierarchical Graph Representation Learning with Differentiable Pooling

Rex Ying, Jiaxuan You, Christopher Morris, Xiang Ren, William L. Hamilton, Jure Leskovec

1806.08804

Unsupervised Domain Adaptation by Backpropagation

Yaroslav Ganin, Victor Lempitsky

1409.7495

Understanding Neural Networks Through Deep Visualization

Jason Yosinski, Jeff Clune, Anh Nguyen, Thomas Fuchs, Hod Lipson

1506.06579

Towards better decoding and language model integration in sequence to sequence models

Jan Chorowski, Navdeep Jaitly

1612.02695

Learning Features by Watching Objects Move

Deepak Pathak, Ross Girshick, Piotr Dollár, Trevor Darrell, Bharath Hariharan

1612.06370

Striving for Simplicity: The All Convolutional Net

Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, Martin Riedmiller

1412.6806

On Learning to Think: Algorithmic Information Theory for Novel Combinations of Reinforcement Learning Controllers and Recurrent Neural World Models

Juergen Schmidhuber

1511.09249

Optimal approximation of continuous functions by very deep ReLU networks

Dmitry Yarotsky

1802.03620

Deep Quaternion Networks

Chase Gaudet, Anthony Maida

1712.04604

Eignets for function approximation on manifolds

H. N. Mhaskar

0909.5000

Natasha 2: Faster Non-Convex Optimization Than SGD

Zeyuan Allen-Zhu

1708.08694

Generative Adversarial Text to Image Synthesis

Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, Honglak Lee

1605.05396

Virtual Worlds as Proxy for Multi-Object Tracking Analysis

Adrien Gaidon, Qiao Wang, Yohann Cabon, Eleonora Vig

1605.06457

End-to-End Learning of Geometry and Context for Deep Stereo Regression

Alex Kendall, Hayk Martirosyan, Saumitro Dasgupta, Peter Henry, Ryan Kennedy, Abraham Bachrach, Adam Bry

1703.04309

Visualizing and Understanding Recurrent Networks

Andrej Karpathy, Justin Johnson, Li Fei-Fei

1506.02078

Finding Approximate Local Minima Faster than Gradient Descent

Naman Agarwal, Zeyuan Allen-Zhu, Brian Bullins, Elad Hazan, Tengyu Ma

1611.01146

Discriminative Unsupervised Feature Learning with Exemplar Convolutional Neural Networks

Alexey Dosovitskiy, Philipp Fischer, Jost Tobias Springenberg, Martin Riedmiller, Thomas Brox

1406.6909

RETURNN as a Generic Flexible Neural Toolkit with Application to Translation and Speech Recognition

Albert Zeyer, Tamer Alkhouli, Hermann Ney

1805.05225

Improving speech recognition by revising gated recurrent units

Mirco Ravanelli, Philemon Brakel, Maurizio Omologo, Yoshua Bengio

1710.00641

QMDP-Net: Deep Learning for Planning under Partial Observability

Peter Karkus, David Hsu, Wee Sun Lee

1703.06692

Learning to Perform Physics Experiments via Deep Reinforcement Learning

Misha Denil, Pulkit Agrawal, Tejas D Kulkarni, Tom Erez, Peter Battaglia, Nando de Freitas

1611.01843

The Hessian Penalty: A Weak Prior for Unsupervised Disentanglement

William Peebles, John Peebles, Jun-Yan Zhu, Alexei Efros, Antonio Torralba

2008.10599

Measuring Neural Net Robustness with Constraints

Osbert Bastani, Yani Ioannou, Leonidas Lampropoulos, Dimitrios Vytiniotis, Aditya Nori, Antonio Criminisi

1605.07262

Training Generative Adversarial Networks with Limited Data

Tero Karras, Miika Aittala, Janne Hellsten, Samuli Laine, Jaakko Lehtinen, Timo Aila

2006.06676

Unsupervised Pretraining for Sequence to Sequence Learning

Prajit Ramachandran, Peter J. Liu, Quoc V. Le

1611.02683

Adaptive Computation Time for Recurrent Neural Networks

Alex Graves

1603.08983

Exploring the Space of Adversarial Images

Pedro Tabacof, Eduardo Valle

1510.05328

Adversarial Manipulation of Deep Representations

Sara Sabour, Yanshuai Cao, Fartash Faghri, David J. Fleet

1511.05122

Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling

Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, Yoshua Bengio

1412.3555

Coupled Oscillatory Recurrent Neural Network (coRNN): An accurate and (gradient) stable architecture for learning long time dependencies

T. Konstantin Rusch, Siddhartha Mishra

2010.00951

Bayesian SegNet: Model Uncertainty in Deep Convolutional Encoder-Decoder Architectures for Scene Understanding

Alex Kendall, Vijay Badrinarayanan, Roberto Cipolla

1511.02680

Minimal Gated Unit for Recurrent Neural Networks

Guo-Bing Zhou, Jianxin Wu, Chen-Lin Zhang, Zhi-Hua Zhou

1603.09420

Value Iteration Networks

Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, Pieter Abbeel

1602.02867

Transformation Properties of Learned Visual Representations

Taco S. Cohen, Max Welling

1412.7659

Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy

Asit Mishra, Debbie Marr

1711.05852

Mean Field Residual Networks: On the Edge of Chaos

Greg Yang, Samuel S. Schoenholz

1712.08969

Learning Structured Sparsity in Deep Neural Networks

Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, Hai Li

1608.03665

Interpretable Convolutional Filters with SincNet

Mirco Ravanelli, Yoshua Bengio

1811.09725

The PyTorch-Kaldi Speech Recognition Toolkit

Mirco Ravanelli, Titouan Parcollet, Yoshua Bengio

1811.07453

Adversarial Generation of Natural Language

Sai Rajeswar, Sandeep Subramanian, Francis Dutil, Christopher Pal, Aaron Courville

1705.10929

Self-labelling via simultaneous clustering and representation learning

Yuki Markus Asano, Christian Rupprecht, Andrea Vedaldi

1911.05371

Mode Regularized Generative Adversarial Networks

Tong Che, Yanran Li, Athul Paul Jacob, Yoshua Bengio, Wenjie Li

1612.02136

Consistency by Agreement in Zero-shot Neural Machine Translation

Maruan Al-Shedivat, Ankur P. Parikh

1904.02338

Data Augmentation Generative Adversarial Networks

Antreas Antoniou, Amos Storkey, Harrison Edwards

1711.04340

Deep CORAL: Correlation Alignment for Deep Domain Adaptation

Baochen Sun, Kate Saenko

1607.01719

Can recurrent neural networks warp time?

Corentin Tallec, Yann Ollivier

1804.11188

TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning

Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, Hai Li

1705.07878

The case for 4-bit precision: k-bit Inference Scaling Laws

Tim Dettmers, Luke Zettlemoyer

2212.09720

Multimodal Deep Learning for Robust RGB-D Object Recognition

Andreas Eitel, Jost Tobias Springenberg, Luciano Spinello, Martin Riedmiller, Wolfram Burgard

1507.06821

Dynamic Network Surgery for Efficient DNNs

Yiwen Guo, Anbang Yao, Yurong Chen

1608.04493

Embracing data abundance: BookTest Dataset for Reading Comprehension

Ondrej Bajgar, Rudolf Kadlec, Jan Kleindienst

1610.00956

Multistep Distillation of Diffusion Models via Moment Matching

Tim Salimans, Thomas Mensink, Jonathan Heek, Emiel Hoogeboom

2406.04103

Hyperprior Induced Unsupervised Disentanglement of Latent Representations

Abdul Fatir Ansari, Harold Soh

1809.04497

Latent Intention Dialogue Models

Tsung-Hsien Wen, Yishu Miao, Phil Blunsom, Steve Young

1705.10229

The Goldilocks zone: Towards better understanding of neural network loss landscapes

Stanislav Fort, Adam Scherlis

1807.02581

Neon2: Finding Local Minima via First-Order Oracles

Zeyuan Allen-Zhu, Yuanzhi Li

1711.06673

A Convergence Analysis of Gradient Descent for Deep Linear Neural Networks

Sanjeev Arora, Nadav Cohen, Noah Golowich, Wei Hu

1810.02281

Lifted Relational Neural Networks

Gustav Sourek, Vojtech Aschenbrenner, Filip Zelezny, Ondrej Kuzelka

1508.05128

Light Gated Recurrent Units for Speech Recognition

Mirco Ravanelli, Philemon Brakel, Maurizio Omologo, Yoshua Bengio

1803.10225

Convolutional Neural Networks Applied to House Numbers Digit Classification

Pierre Sermanet, Soumith Chintala, Yann LeCun

1204.3968

Detecting and Correcting for Label Shift with Black Box Predictors

Zachary C. Lipton, Yu-Xiang Wang, Alex Smola

1802.03916

Swapout: Learning an ensemble of deep architectures

Saurabh Singh, Derek Hoiem, David Forsyth

1605.06465

Learning Important Features Through Propagating Activation Differences

Avanti Shrikumar, Peyton Greenside, Anshul Kundaje

1704.02685

DiCE: The Infinitely Differentiable Monte-Carlo Estimator

Jakob Foerster, Gregory Farquhar, Maruan Al-Shedivat, Tim Rocktäschel, Eric P. Xing, Shimon Whiteson

1802.05098

The Mechanics of n-Player Differentiable Games

David Balduzzi, Sebastien Racaniere, James Martens, Jakob Foerster, Karl Tuyls, Thore Graepel

1802.05642

Are Disentangled Representations Helpful for Abstract Visual Reasoning?

Sjoerd van Steenkiste, Francesco Locatello, Jürgen Schmidhuber, Olivier Bachem

1905.12506

Data-Driven Sparse Structure Selection for Deep Neural Networks

Zehao Huang, Naiyan Wang

1707.01213

What to talk about and how? Selective Generation using LSTMs with Coarse-to-Fine Alignment

Hongyuan Mei, Mohit Bansal, Matthew R. Walter

1509.00838

PDE-Net: Learning PDEs from Data

Zichao Long, Yiping Lu, Xianzhong Ma, Bin Dong

1710.09668

A Clockwork RNN

Jan Koutník, Klaus Greff, Faustino Gomez, Jürgen Schmidhuber

1402.3511

MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems

Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, Zheng Zhang

1512.01274

Deep SimNets

Nadav Cohen, Or Sharir, Amnon Shashua

1506.03059

Deep Learning in Neural Networks: An Overview

Juergen Schmidhuber

1404.7828

Spatio-Temporal LSTM with Trust Gates for 3D Human Action Recognition

Jun Liu, Amir Shahroudy, Dong Xu, Gang Wang

1607.07043

Full-Capacity Unitary Recurrent Neural Networks

Scott Wisdom, Thomas Powers, John R. Hershey, Jonathan Le Roux, Les Atlas

1611.00035

Return of Frustratingly Easy Domain Adaptation

Baochen Sun, Jiashi Feng, Kate Saenko

1511.05547

Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics

Norman Di Palo, Edward Johns

2403.19578

Match-SRNN: Modeling the Recursive Matching Structure with Spatial RNN

Shengxian Wan, Yanyan Lan, Jun Xu, Jiafeng Guo, Liang Pang, Xueqi Cheng

1604.04378

Dropout improves Recurrent Neural Networks for Handwriting Recognition

Vu Pham, Théodore Bluche, Christopher Kermorvant, Jérôme Louradour

1312.4569

Hamiltonian Neural Networks

Sam Greydanus, Misko Dzamba, Jason Yosinski

1906.01563

Supervised Speech Separation Based on Deep Learning: An Overview

DeLiang Wang, Jitong Chen

1708.07524

Collaborative Deep Learning for Recommender Systems

Hao Wang, Naiyan Wang, Dit-Yan Yeung

1409.2944

A Deep Architecture for Semantic Matching with Multiple Positional Sentence Representations

Shengxian Wan, Yanyan Lan, Jiafeng Guo, Jun Xu, Liang Pang, Xueqi Cheng

1511.08277

Recurrent Human Pose Estimation

Vasileios Belagiannis, Andrew Zisserman

1605.02914

Filtering Variational Objectives

Chris J. Maddison, Dieterich Lawson, George Tucker, Nicolas Heess, Mohammad Norouzi, Andriy Mnih, Arnaud Doucet, Yee Whye Teh

1705.09279

Residual Networks Behave Like Ensembles of Relatively Shallow Networks

Andreas Veit, Michael Wilber, Serge Belongie

1605.06431

Generating Long Videos of Dynamic Scenes

Tim Brooks, Janne Hellsten, Miika Aittala, Ting-Chun Wang, Timo Aila, Jaakko Lehtinen, Ming-Yu Liu, Alexei A. Efros, Tero Karras

2206.03429

Analysis Methods in Neural Language Processing: A Survey

Yonatan Belinkov, James Glass

1812.08951

The Role of ImageNet Classes in Fréchet Inception Distance

Tuomas Kynkäänniemi, Tero Karras, Miika Aittala, Timo Aila, Jaakko Lehtinen

2203.06026

Scalable Adaptive Computation for Iterative Generation

Allan Jabri, David Fleet, Ting Chen

2212.11972

Move Evaluation in Go Using Deep Convolutional Neural Networks

Chris J. Maddison, Aja Huang, Ilya Sutskever, David Silver

1412.6564

Stacked Generative Adversarial Networks

Xun Huang, Yixuan Li, Omid Poursaeed, John Hopcroft, Serge Belongie

1612.04357

First-Pass Large Vocabulary Continuous Speech Recognition using Bi-Directional Recurrent DNNs

Awni Y. Hannun, Andrew L. Maas, Daniel Jurafsky, Andrew Y. Ng

1408.2873

General-Purpose In-Context Learning by Meta-Learning Transformers

Louis Kirsch, James Harrison, Jascha Sohl-Dickstein, Luke Metz

2212.04458

Sequence Transduction with Recurrent Neural Networks

Alex Graves

1211.3711

Sequence-Level Knowledge Distillation

Yoon Kim, Alexander M. Rush

1606.07947

Generating Sentences by Editing Prototypes

Kelvin Guu, Tatsunori B. Hashimoto, Yonatan Oren, Percy Liang

1709.08878

Generating Factoid Questions With Recurrent Neural Networks: The 30M Factoid Question-Answer Corpus

Iulian Vlad Serban, Alberto García-Durán, Caglar Gulcehre, Sungjin Ahn, Sarath Chandar, Aaron Courville, Yoshua Bengio

1603.06807

Generative Choreography using Deep Learning

Luka Crnkovic-Friis, Louise Crnkovic-Friis

1605.06921

Towards mental time travel: a hierarchical memory for reinforcement learning agents

Andrew Kyle Lampinen, Stephanie C. Y. Chan, Andrea Banino, Felix Hill

2105.14039

How Does Batch Normalization Help Optimization?

Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, Aleksander Madry

1805.11604

Classifying Relations by Ranking with Convolutional Neural Networks

Cicero Nogueira dos Santos, Bing Xiang, Bowen Zhou

1504.06580

Enhanced Convolutional Neural Tangent Kernels

Zhiyuan Li, Ruosong Wang, Dingli Yu, Simon S. Du, Wei Hu, Ruslan Salakhutdinov, Sanjeev Arora

1911.00809

The Predictron: End-To-End Learning and Planning

David Silver, Hado van Hasselt, Matteo Hessel, Tom Schaul, Arthur Guez, Tim Harley, Gabriel Dulac-Arnold, David Reichert, Neil Rabinowitz, Andre Barreto, Thomas Degris

1612.08810

On orthogonality and learning recurrent networks with long term dependencies

Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, Chris Pal

1702.00071

A New Training Pipeline for an Improved Neural Transducer

Albert Zeyer, André Merboldt, Ralf Schlüter, Hermann Ney

2005.09319

Overcoming Exploration in Reinforcement Learning with Demonstrations

Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, Pieter Abbeel

1709.10089

Memory Visualization for Gated Recurrent Neural Networks in Speech Recognition

Zhiyuan Tang, Ying Shi, Dong Wang, Yang Feng, Shiyue Zhang

1609.08789

FBNetV2: Differentiable Neural Architecture Search for Spatial and Channel Dimensions

Alvin Wan, Xiaoliang Dai, Peizhao Zhang, Zijian He, Yuandong Tian, Saining Xie, Bichen Wu, Matthew Yu, Tao Xu, Kan Chen, Peter Vajda, Joseph E. Gonzalez

2004.05565

Understanding image representations by measuring their equivariance and equivalence

Karel Lenc, Andrea Vedaldi

1411.5908

The Lottery Tickets Hypothesis for Supervised and Self-supervised Pre-training in Computer Vision Models

Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Michael Carbin, Zhangyang Wang

2012.06908

Every Model Learned by Gradient Descent Is Approximately a Kernel Machine

Pedro Domingos

2012.00152

Tunable Efficient Unitary Neural Networks (EUNN) and their application to RNNs

Li Jing, Yichen Shen, Tena Dubček, John Peurifoy, Scott Skirlo, Yann LeCun, Max Tegmark, Marin Soljačić

1612.05231

On the Binding Problem in Artificial Neural Networks

Klaus Greff, Sjoerd van Steenkiste, Jürgen Schmidhuber

2012.05208

Angle-based Search Space Shrinking for Neural Architecture Search

Yiming Hu, Yuding Liang, Zichao Guo, Ruosi Wan, Xiangyu Zhang, Yichen Wei, Qingyi Gu, Jian Sun

2004.13431

Single Headed Attention RNN: Stop Thinking With Your Head

Stephen Merity

1911.11423

Named Entity Recognition with Bidirectional LSTM-CNNs

Jason P. C. Chiu, Eric Nichols

1511.08308

Feature Learning in Infinite-Width Neural Networks

Greg Yang, Edward J. Hu

2011.14522

"Zero-Shot" Super-Resolution using Deep Internal Learning

Assaf Shocher, Nadav Cohen, Michal Irani

1712.06087

Vote3Deep: Fast Object Detection in 3D Point Clouds Using Efficient Convolutional Neural Networks

Martin Engelcke, Dushyant Rao, Dominic Zeng Wang, Chi Hay Tong, Ingmar Posner

1609.06666

Depth-Width Tradeoffs in Approximating Natural Functions with Neural Networks

Itay Safran, Ohad Shamir

1610.09887

Feature Purification: How Adversarial Training Performs Robust Deep Learning

Zeyuan Allen-Zhu, Yuanzhi Li

2005.10190

Learning Explanatory Rules from Noisy Data

Richard Evans, Edward Grefenstette

1711.04574

Transfer Learning for Named-Entity Recognition with Neural Networks

Ji Young Lee, Franck Dernoncourt, Peter Szolovits

1705.06273

Latent Constraints: Learning to Generate Conditionally from Unconditional Generative Models

Jesse Engel, Matthew Hoffman, Adam Roberts

1711.05772

Alias-Free Generative Adversarial Networks

Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, Timo Aila

2106.12423

AdaBits: Neural Network Quantization with Adaptive Bit-Widths

Qing Jin, Linjie Yang, Zhenyu Liao

1912.09666

Gram-CTC: Automatic Unit Selection and Target Decomposition for Sequence Labelling

Hairong Liu, Zhenyao Zhu, Xiangang Li, Sanjeev Satheesh

1703.00096

Deep Neural Networks with Random Gaussian Weights: A Universal Classification Strategy?

Raja Giryes, Guillermo Sapiro, Alex M. Bronstein

1504.08291

Adversarial Perturbations Against Deep Neural Networks for Malware Classification

Kathrin Grosse, Nicolas Papernot, Praveen Manoharan, Michael Backes, Patrick McDaniel

1606.04435

The Surprising Simplicity of the Early-Time Learning Dynamics of Neural Networks

Wei Hu, Lechao Xiao, Ben Adlam, Jeffrey Pennington

2006.14599

A Deep Reinforcement Learning Chatbot

Iulian V. Serban, Chinnadhurai Sankar, Mathieu Germain, Saizheng Zhang, Zhouhan Lin, Sandeep Subramanian, Taesup Kim, Michael Pieper, Sarath Chandar, Nan Rosemary Ke, Sai Rajeshwar, Alexandre de Brebisson, Jose M. R. Sotelo, Dendi Suhubdy, Vincent Michalski, Alexandre Nguyen, Joelle Pineau, Yoshua Bengio

1709.02349

On Exact Computation with an Infinitely Wide Neural Net

Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, Ruosong Wang

1904.11955

Explaining Recurrent Neural Network Predictions in Sentiment Analysis

Leila Arras, Grégoire Montavon, Klaus-Robert Müller, Wojciech Samek

1706.07206

DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients

Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, Yuheng Zou

1606.06160

Structured Attention Networks

Yoon Kim, Carl Denton, Luong Hoang, Alexander M. Rush

1702.00887

What Can ResNet Learn Efficiently, Going Beyond Kernels?

Zeyuan Allen-Zhu, Yuanzhi Li

1905.10337

Temporal Ensembling for Semi-Supervised Learning

Samuli Laine, Timo Aila

1610.02242

Semi-supervised Multitask Learning for Sequence Labeling

Marek Rei

1704.07156

Evolutionary Generative Adversarial Networks

Chaoyue Wang, Chang Xu, Xin Yao, Dacheng Tao

1803.00657

Searching for Activation Functions

Prajit Ramachandran, Barret Zoph, Quoc V. Le

1710.05941

Insights on representational similarity in neural networks with canonical correlation

Ari S. Morcos, Maithra Raghu, Samy Bengio

1806.05759

Convergent Learning: Do different neural networks learn the same representations?

Yixuan Li, Jason Yosinski, Jeff Clune, Hod Lipson, John Hopcroft

1511.07543

An Analysis of Neural Language Modeling at Multiple Scales

Stephen Merity, Nitish Shirish Keskar, Richard Socher

1803.08240

End-to-End Attention-based Large Vocabulary Speech Recognition

Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, Yoshua Bengio

1508.04395

Deep Unfolding: Model-Based Inspiration of Novel Deep Architectures

John R. Hershey, Jonathan Le Roux, Felix Weninger

1409.2574

Structured Pruning of Deep Convolutional Neural Networks

Sajid Anwar, Kyuyeon Hwang, Wonyong Sung

1512.08571

Explaining How a Deep Neural Network Trained with End-to-End Learning Steers a Car

Mariusz Bojarski, Philip Yeres, Anna Choromanska, Krzysztof Choromanski, Bernhard Firner, Lawrence Jackel, Urs Muller

1704.07911

No bad local minima: Data independent training error guarantees for multilayer neural networks

Daniel Soudry, Yair Carmon

1605.08361

Domain-Adversarial Neural Networks

Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand

1412.4446

Inverting Visual Representations with Convolutional Networks

Alexey Dosovitskiy, Thomas Brox

1506.02753

Reasoning About Pragmatics with Neural Listeners and Speakers

Jacob Andreas, Dan Klein

1604.00562

Generative Deep Neural Networks for Dialogue: A Short Review

Iulian Vlad Serban, Ryan Lowe, Laurent Charlin, Joelle Pineau

1611.06216

Neural Networks with Few Multiplications

Zhouhan Lin, Matthieu Courbariaux, Roland Memisevic, Yoshua Bengio

1510.03009

Optimizing Neural Networks with Kronecker-factored Approximate Curvature

James Martens, Roger Grosse

1503.05671

Tagger: Deep Unsupervised Perceptual Grouping

Klaus Greff, Antti Rasmus, Mathias Berglund, Tele Hotloo Hao, Jürgen Schmidhuber, Harri Valpola

1606.06724

Exploring the Landscape of Spatial Robustness

Logan Engstrom, Brandon Tran, Dimitris Tsipras, Ludwig Schmidt, Aleksander Madry

1712.02779

Elucidating the Design Space of Diffusion-Based Generative Models

Tero Karras, Miika Aittala, Timo Aila, Samuli Laine

2206.00364

Deep Reinforcement Learning with Distributional Semantic Rewards for Abstractive Summarization

Siyao Li, Deren Lei, Pengda Qin, William Yang Wang

1909.00141

Optimal Regularization Can Mitigate Double Descent

Preetum Nakkiran, Prayaag Venkat, Sham Kakade, Tengyu Ma

2003.01897

Learning Deep Object Detectors from 3D Models

Xingchao Peng, Baochen Sun, Karim Ali, Kate Saenko

1412.7122

Tensor Programs I: Wide Feedforward or Recurrent Neural Networks of Any Architecture are Gaussian Processes

Greg Yang

1910.12478

Semantic3D.net: A new Large-scale Point Cloud Classification Benchmark

Timo Hackel, Nikolay Savinov, Lubor Ladicky, Jan D. Wegner, Konrad Schindler, Marc Pollefeys

1704.03847

AtomNet: A Deep Convolutional Neural Network for Bioactivity Prediction in Structure-based Drug Discovery

Izhar Wallach, Michael Dzamba, Abraham Heifets

1510.02855

Taskonomy: Disentangling Task Transfer Learning

Amir Zamir, Alexander Sax, William Shen, Leonidas Guibas, Jitendra Malik, Silvio Savarese

1804.08328

A Genetic Programming Approach to Designing Convolutional Neural Network Architectures

Masanori Suganuma, Shinichi Shirakawa, Tomoharu Nagao

1704.00764

Deep Directed Generative Autoencoders

Sherjil Ozair, Yoshua Bengio

1410.0630

Trainable Frontend For Robust and Far-Field Keyword Spotting

Yuxuan Wang, Pascal Getreuer, Thad Hughes, Richard F. Lyon, Rif A. Saurous

1607.05666

Modular Duality in Deep Learning

Jeremy Bernstein, Laker Newhouse

2410.21265

Very Deep Convolutional Neural Networks for Raw Waveforms

Wei Dai, Chia Dai, Shuhui Qu, Juncheng Li, Samarjit Das

1610.00087

Resiliency of Deep Neural Networks under Quantization

Wonyong Sung, Sungho Shin, Kyuyeon Hwang

1511.06488

Online Batch Selection for Faster Training of Neural Networks

Ilya Loshchilov, Frank Hutter

1511.06343

A Focused Dynamic Attention Model for Visual Question Answering

Ilija Ilievski, Shuicheng Yan, Jiashi Feng

1604.01485

Neural Programmer-Interpreters

Scott Reed, Nando de Freitas

1511.06279

Bias Correction of Learned Generative Models using Likelihood-Free Importance Weighting

Aditya Grover, Jiaming Song, Alekh Agarwal, Kenneth Tran, Ashish Kapoor, Eric Horvitz, Stefano Ermon

1906.09531

Convolutional Neural Networks using Logarithmic Data Representation

Daisuke Miyashita, Edward H. Lee, Boris Murmann

1603.01025

Hindsight Experience Replay

Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, Wojciech Zaremba

1707.01495

Self-supervised deep convolutional neural network for chest X-ray classification

Matej Gazda, Jakub Gazda, Jan Plavka, Peter Drotar

2103.03055

Stochastic Optimization of Sorting Networks via Continuous Relaxations

Aditya Grover, Eric Wang, Aaron Zweig, Stefano Ermon

1903.08850

On the saddle point problem for non-convex optimization

Razvan Pascanu, Yann N. Dauphin, Surya Ganguli, Yoshua Bengio

1405.4604

DialogWAE: Multimodal Response Generation with Conditional Wasserstein Auto-Encoder

Xiaodong Gu, Kyunghyun Cho, Jung-Woo Ha, Sunghun Kim

1805.12352

Exact solutions to the nonlinear dynamics of learning in deep linear neural networks

Andrew M. Saxe, James L. McClelland, Surya Ganguli

1312.6120

Directional convergence and alignment in deep learning

Ziwei Ji, Matus Telgarsky

2006.06657

Flow++: Improving Flow-Based Generative Models with Variational Dequantization and Architecture Design

Jonathan Ho, Xi Chen, Aravind Srinivas, Yan Duan, Pieter Abbeel

1902.00275

Universal approximations of invariant maps by neural networks

Dmitry Yarotsky

1804.10306

Scalable Training of Artificial Neural Networks with Adaptive Sparse Connectivity inspired by Network Science

Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong H. Nguyen, Madeleine Gibescu, Antonio Liotta

1707.04780

Generalizing Pooling Functions in Convolutional Neural Networks: Mixed, Gated, and Tree

Chen-Yu Lee, Patrick W. Gallagher, Zhuowen Tu

1509.08985

Imitation from Observation: Learning to Imitate Behaviors from Raw Video via Context Translation

YuXuan Liu, Abhishek Gupta, Pieter Abbeel, Sergey Levine

1707.03374

Recurrent Neural Network Training with Dark Knowledge Transfer

Zhiyuan Tang, Dong Wang, Zhiyong Zhang

1505.04630

Review of Deep Learning

Rong Zhang, Weiping Li, Tong Mo

1804.01653

Gated Feedback Recurrent Neural Networks

Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, Yoshua Bengio

1502.02367

Streaming PCA: Matching Matrix Bernstein and Near-Optimal Finite Sample Guarantees for Oja's Algorithm

Prateek Jain, Chi Jin, Sham M. Kakade, Praneeth Netrapalli, Aaron Sidford

1602.06929

On the Limitations of Representing Functions on Sets

Edward Wagstaff, Fabian B. Fuchs, Martin Engelcke, Ingmar Posner, Michael Osborne

1901.09006

Unity: A General Platform for Intelligent Agents

Arthur Juliani, Vincent-Pierre Berges, Ervin Teng, Andrew Cohen, Jonathan Harper, Chris Elion, Chris Goy, Yuan Gao, Hunter Henry, Marwan Mattar, Danny Lange

1809.02627

Junction Tree Variational Autoencoder for Molecular Graph Generation

Wengong Jin, Regina Barzilay, Tommi Jaakkola

1802.04364

Large Language Model Guided Tree-of-Thought

Jieyi Long

2305.08291

A Joint Model for Question Answering and Question Generation

Tong Wang, Xingdi Yuan, Adam Trischler

1706.01450

The Neural Noisy Channel

Lei Yu, Phil Blunsom, Chris Dyer, Edward Grefenstette, Tomas Kocisky

1611.02554

ReasoNet: Learning to Stop Reading in Machine Comprehension

Yelong Shen, Po-Sen Huang, Jianfeng Gao, Weizhu Chen

1609.05284

From neural PCA to deep unsupervised learning

Harri Valpola

1411.7783

Character-based Neural Machine Translation

Marta R. Costa-Jussà, José A. R. Fonollosa

1603.00810

Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning

William Lotter, Gabriel Kreiman, David Cox

1605.08104

DeepID-Net: Deformable Deep Convolutional Neural Networks for Object Detection

Wanli Ouyang, Xiaogang Wang, Xingyu Zeng, Shi Qiu, Ping Luo, Yonglong Tian, Hongsheng Li, Shuo Yang, Zhe Wang, Chen-Change Loy, Xiaoou Tang

1412.5661

Measuring the Intrinsic Dimension of Objective Landscapes

Chunyuan Li, Heerad Farkhoor, Rosanne Liu, Jason Yosinski

1804.08838

Why does deep and cheap learning work so well?

Henry W. Lin, Max Tegmark, David Rolnick

1608.08225

Differentiable Game Mechanics

Alistair Letcher, David Balduzzi, Sebastien Racaniere, James Martens, Jakob Foerster, Karl Tuyls, Thore Graepel

1905.04926

On the Power and Limitations of Random Features for Understanding Neural Networks

Gilad Yehudai, Ohad Shamir

1904.00687

A guide to convolution arithmetic for deep learning

Vincent Dumoulin, Francesco Visin

1603.07285

Gradient Descent Maximizes the Margin of Homogeneous Neural Networks

Kaifeng Lyu, Jian Li

1906.05890

Object Detectors Emerge in Deep Scene CNNs

Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, Antonio Torralba

1412.6856

Bitwise Neural Networks

Minje Kim, Paris Smaragdis

1601.06071

Deep neural networks are robust to weight binarization and other non-linear distortions

Paul Merolla, Rathinakumar Appuswamy, John Arthur, Steve K. Esser, Dharmendra Modha

1606.01981

Do Convnets Learn Correspondence?

Jonathan Long, Ning Zhang, Trevor Darrell

1411.1091

Deep Fried Convnets

Zichao Yang, Marcin Moczulski, Misha Denil, Nando de Freitas, Alex Smola, Le Song, Ziyu Wang

1412.7149

Meta-Learning and Universality: Deep Representations and Gradient Descent can Approximate any Learning Algorithm

Chelsea Finn, Sergey Levine

1710.11622

Can Neural Networks Understand Logical Entailment?

Richard Evans, David Saxton, David Amos, Pushmeet Kohli, Edward Grefenstette

1802.08535

Dataset and Neural Recurrent Sequence Labeling Model for Open-Domain Factoid Question Answering

Peng Li, Wei Li, Zhengyan He, Xuguang Wang, Ying Cao, Jie Zhou, Wei Xu

1607.06275

Do Deep Nets Really Need to be Deep?

Lei Jimmy Ba, Rich Caruana

1312.6184

Benchmarking Neural Network Robustness to Common Corruptions and Surface Variations

Dan Hendrycks, Thomas G. Dietterich

1807.01697

High-Performance Neural Networks for Visual Object Classification

Dan C. Cireşan, Ueli Meier, Jonathan Masci, Luca M. Gambardella, Jürgen Schmidhuber

1102.0183

Relational Neural Expectation Maximization: Unsupervised Discovery of Objects and their Interactions

Sjoerd van Steenkiste, Michael Chang, Klaus Greff, Jürgen Schmidhuber

1802.10353

PathNet: Evolution Channels Gradient Descent in Super Neural Networks

Chrisantha Fernando, Dylan Banarse, Charles Blundell, Yori Zwols, David Ha, Andrei A. Rusu, Alexander Pritzel, Daan Wierstra

1701.08734

CNN-RNN: A Unified Framework for Multi-label Image Classification

Jiang Wang, Yi Yang, Junhua Mao, Zhiheng Huang, Chang Huang, Wei Xu

1604.04573

Making Neural QA as Simple as Possible but not Simpler

Dirk Weissenborn, Georg Wiese, Laura Seiffe

1703.04816

Automated Curriculum Learning for Neural Networks

Alex Graves, Marc G. Bellemare, Jacob Menick, Remi Munos, Koray Kavukcuoglu

1704.03003

Jointly Learning to Label Sentences and Tokens

Marek Rei, Anders Søgaard

1811.05949

Compositional Obverter Communication Learning From Raw Visual Input

Edward Choi, Angeliki Lazaridou, Nando de Freitas

1804.02341

Gated Orthogonal Recurrent Units: On Learning to Forget

Li Jing, Caglar Gulcehre, John Peurifoy, Yichen Shen, Max Tegmark, Marin Soljačić, Yoshua Bengio

1706.02761

BinaryConnect: Training Deep Neural Networks with binary weights during propagations

Matthieu Courbariaux, Yoshua Bengio, Jean-Pierre David

1511.00363

Compressing Deep Convolutional Networks using Vector Quantization

Yunchao Gong, Liu Liu, Ming Yang, Lubomir Bourdev

1412.6115

Learning model-based planning from scratch

Razvan Pascanu, Yujia Li, Oriol Vinyals, Nicolas Heess, Lars Buesing, Sebastien Racanière, David Reichert, Théophane Weber, Daan Wierstra, Peter Battaglia

1707.06170

A Supervised Approach to Extractive Summarisation of Scientific Papers

Ed Collins, Isabelle Augenstein, Sebastian Riedel

1706.03946

Tensor Programs III: Neural Matrix Laws

Greg Yang

2009.10685

Random Walk Initialization for Training Very Deep Feedforward Networks

David Sussillo, L. F. Abbott

1412.6558

Highway Long Short-Term Memory RNNs for Distant Speech Recognition

Yu Zhang, Guoguo Chen, Dong Yu, Kaisheng Yao, Sanjeev Khudanpur, James Glass

1510.08983

Density estimation using Real NVP

Laurent Dinh, Jascha Sohl-Dickstein, Samy Bengio

1605.08803

Learning the Number of Neurons in Deep Networks

Jose M Alvarez, Mathieu Salzmann

1611.06321

Building Program Vector Representations for Deep Learning

Lili Mou, Ge Li, Yuxuan Liu, Hao Peng, Zhi Jin, Yan Xu, Lu Zhang

1409.3358

Massively Parallel Methods for Deep Reinforcement Learning

Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, Shane Legg, Volodymyr Mnih, Koray Kavukcuoglu, David Silver

1507.04296

The loss surface of deep and wide neural networks

Quynh Nguyen, Matthias Hein

1704.08045

ChamNet: Towards Efficient Network Design through Platform-Aware Model Adaptation

Xiaoliang Dai, Peizhao Zhang, Bichen Wu, Hongxu Yin, Fei Sun, Yanghan Wang, Marat Dukhan, Yunqing Hu, Yiming Wu, Yangqing Jia, Peter Vajda, Matt Uyttendaele, Niraj K. Jha

1812.08934

Feed-Forward Networks with Attention Can Solve Some Long-Term Memory Problems

Colin Raffel, Daniel P. W. Ellis

1512.08756

Automatic Gradient Descent: Deep Learning without Hyperparameters

Jeremy Bernstein, Chris Mingard, Kevin Huang, Navid Azizan, Yisong Yue

2304.05187

Grid Long Short-Term Memory

Nal Kalchbrenner, Ivo Danihelka, Alex Graves

1507.01526

An exact mapping between the Variational Renormalization Group and Deep Learning

Pankaj Mehta, David J. Schwab

1410.3831

A Comprehensive Study of Deep Bidirectional LSTM RNNs for Acoustic Modeling in Speech Recognition

Albert Zeyer, Patrick Doetsch, Paul Voigtlaender, Ralf Schlüter, Hermann Ney

1606.06871

Relevance-based Word Embedding

Hamed Zamani, W. Bruce Croft

1705.03556

Variational Walkback: Learning a Transition Operator as a Stochastic Recurrent Net

Anirudh Goyal, Nan Rosemary Ke, Surya Ganguli, Yoshua Bengio

1711.02282

In Defense of the Triplet Loss for Person Re-Identification

Alexander Hermans, Lucas Beyer, Bastian Leibe

1703.07737

Very Deep Multilingual Convolutional Neural Networks for LVCSR

Tom Sercu, Christian Puhrsch, Brian Kingsbury, Yann LeCun

1509.08967

Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping

James Martens, Andy Ballard, Guillaume Desjardins, Grzegorz Swirszcz, Valentin Dalibard, Jascha Sohl-Dickstein, Samuel S. Schoenholz

2110.01765

Unitary Evolution Recurrent Neural Networks

Martin Arjovsky, Amar Shah, Yoshua Bengio

1511.06464

Adversarially Robust Generalization Requires More Data

Ludwig Schmidt, Shibani Santurkar, Dimitris Tsipras, Kunal Talwar, Aleksander Mądry

1804.11285

Multi-GPU Training of ConvNets

Omry Yadan, Keith Adams, Yaniv Taigman, Marc'Aurelio Ranzato

1312.5853

Learning Natural Language Inference with LSTM

Shuohang Wang, Jing Jiang

1512.08849

Generic 3D Representation via Pose Estimation and Matching

Amir R. Zamir, Tilman Wekel, Pulkit Argrawal, Colin Weil, Jitendra Malik, Silvio Savarese

1710.08247

Sequence-to-Sequence Learning as Beam-Search Optimization

Sam Wiseman, Alexander M. Rush

1606.02960

On the Compression of Recurrent Neural Networks with an Application to LVCSR acoustic modeling for Embedded Speech Recognition

Rohit Prabhavalkar, Ouais Alsharif, Antoine Bruguier, Ian McGraw

1603.08042

Are Neural Nets Modular? Inspecting Functional Modularity Through Differentiable Weight Masks

Róbert Csordás, Sjoerd van Steenkiste, Jürgen Schmidhuber

2010.02066

AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression Rates

Ning Liu, Xiaolong Ma, Zhiyuan Xu, Yanzhi Wang, Jian Tang, Jieping Ye

1907.03141

Submanifold Sparse Convolutional Networks

Benjamin Graham, Laurens van der Maaten

1706.01307

MarrNet: 3D Shape Reconstruction via 2.5D Sketches

Jiajun Wu, Yifan Wang, Tianfan Xue, Xingyuan Sun, William T Freeman, Joshua B Tenenbaum

1711.03129

Deep Successor Reinforcement Learning

Tejas D. Kulkarni, Ardavan Saeedi, Simanta Gautam, Samuel J. Gershman

1606.02396

Graying the black box: Understanding DQNs

Tom Zahavy, Nir Ben Zrihem, Shie Mannor

1602.02658

Attention for Fine-Grained Categorization

Pierre Sermanet, Andrea Frome, Esteban Real

1412.7054

Learning to Prune Deep Neural Networks via Layer-wise Optimal Brain Surgeon

Xin Dong, Shangyu Chen, Sinno Jialin Pan

1705.07565

Analyzing the Performance of Multilayer Neural Networks for Object Recognition

Pulkit Agrawal, Ross Girshick, Jitendra Malik

1407.1610

Compositional Sequence Labeling Models for Error Detection in Learner Writing

Marek Rei, Helen Yannakoudakis

1607.06153

Attending to Characters in Neural Sequence Labeling Models

Marek Rei, Gamal K. O. Crichton, Sampo Pyysalo

1611.04361

Learning A Physical Long-term Predictor

Sebastien Ehrhardt, Aron Monszpart, Niloy J. Mitra, Andrea Vedaldi

1703.00247

Compressing Convolutional Neural Networks

Wenlin Chen, James T. Wilson, Stephen Tyree, Kilian Q. Weinberger, Yixin Chen

1506.04449

Structural-RNN: Deep Learning on Spatio-Temporal Graphs

Ashesh Jain, Amir R. Zamir, Silvio Savarese, Ashutosh Saxena

1511.05298

A Simple Way to Initialize Recurrent Networks of Rectified Linear Units

Quoc V. Le, Navdeep Jaitly, Geoffrey E. Hinton

1504.00941

Deep Tracking: Seeing Beyond Seeing Using Recurrent Neural Networks

Peter Ondruska, Ingmar Posner

1602.00991

Deep Dyna-Q: Integrating Planning for Task-Completion Dialogue Policy Learning

Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, Kam-Fai Wong, Shang-Yu Su

1801.06176

Evolving Deep Neural Networks

Risto Miikkulainen, Jason Liang, Elliot Meyerson, Aditya Rawal, Dan Fink, Olivier Francon, Bala Raju, Hormoz Shahrzad, Arshak Navruzyan, Nigel Duffy, Babak Hodjat

1703.00548

Affordances from Human Videos as a Versatile Representation for Robotics

Shikhar Bahl, Russell Mendonca, Lili Chen, Unnat Jain, Deepak Pathak

2304.08488

Model Accuracy and Runtime Tradeoff in Distributed Deep Learning:A Systematic Study

Suyog Gupta, Wei Zhang, Fei Wang

1509.04210

Deep learning-based numerical methods for high-dimensional parabolic partial differential equations and backward stochastic differential equations

Weinan E, Jiequn Han, Arnulf Jentzen

1706.04702

Evolving Mario Levels in the Latent Space of a Deep Convolutional Generative Adversarial Network

Vanessa Volz, Jacob Schrum, Jialin Liu, Simon M. Lucas, Adam Smith, Sebastian Risi

1805.00728

PCANet: A Simple Deep Learning Baseline for Image Classification?

Tsung-Han Chan, Kui Jia, Shenghua Gao, Jiwen Lu, Zinan Zeng, Yi Ma

1404.3606

Deep Complex Networks

Chiheb Trabelsi, Olexa Bilaniuk, Ying Zhang, Dmitriy Serdyuk, Sandeep Subramanian, João Felipe Santos, Soroush Mehri, Negar Rostamzadeh, Yoshua Bengio, Christopher J Pal

1705.09792

Implicit Discourse Relation Classification via Multi-Task Neural Networks

Yang Liu, Sujian Li, Xiaodong Zhang, Zhifang Sui

1603.02776

Dissecting Neural ODEs

Stefano Massaroli, Michael Poli, Jinkyoo Park, Atsushi Yamashita, Hajime Asama

2002.08071

Improving Neural Question Generation using Answer Separation

Yanghoon Kim, Hwanhee Lee, Joongbo Shin, Kyomin Jung

1809.02393

Convolutional Neural Fabrics

Shreyas Saxena, Jakob Verbeek

1606.02492

MoDeep: A Deep Learning Framework Using Motion Features for Human Pose Estimation

Arjun Jain, Jonathan Tompson, Yann LeCun, Christoph Bregler

1409.7963

Implicit Regularization in Deep Learning May Not Be Explainable by Norms

Noam Razin, Nadav Cohen

2005.06398

WRPN: Wide Reduced-Precision Networks

Asit Mishra, Eriko Nurvitadhi, Jeffrey J Cook, Debbie Marr

1709.01134

One-Shot Imitation Learning

Yan Duan, Marcin Andrychowicz, Bradly C. Stadie, Jonathan Ho, Jonas Schneider, Ilya Sutskever, Pieter Abbeel, Wojciech Zaremba

1703.07326

Parallel training of DNNs with Natural Gradient and Parameter Averaging

Daniel Povey, Xiaohui Zhang, Sanjeev Khudanpur

1410.7455

Complex Query Answering with Neural Link Predictors

Erik Arakelyan, Daniel Daza, Pasquale Minervini, Michael Cochez

2011.03459

Exploring Speech Enhancement with Generative Adversarial Networks for Robust Speech Recognition

Chris Donahue, Bo Li, Rohit Prabhavalkar

1711.05747

Emergence of grid-like representations by training recurrent neural networks to perform spatial localization

Christopher J. Cueva, Xue-Xin Wei

1803.07770

Interpretable Deep Neural Networks for Single-Trial EEG Classification

Irene Sturm, Sebastian Bach, Wojciech Samek, Klaus-Robert Müller

1604.08201

TUDataset: A collection of benchmark datasets for learning with graphs

Christopher Morris, Nils M. Kriege, Franka Bause, Kristian Kersting, Petra Mutzel, Marion Neumann

2007.08663

Gradient descent with identity initialization efficiently learns positive definite linear transformations by deep residual networks

Peter L. Bartlett, David P. Helmbold, Philip M. Long

1802.06093

YodaNN: An Architecture for Ultra-Low Power Binary-Weight CNN Acceleration

Renzo Andri, Lukas Cavigelli, Davide Rossi, Luca Benini

1606.05487

Path-Normalized Optimization of Recurrent Neural Networks with ReLU Activations

Behnam Neyshabur, Yuhuai Wu, Ruslan Salakhutdinov, Nathan Srebro

1605.07154

A Continuous Relaxation of Beam Search for End-to-end Training of Neural Sequence Models

Kartik Goyal, Graham Neubig, Chris Dyer, Taylor Berg-Kirkpatrick

1708.00111

Snips Voice Platform: an embedded Spoken Language Understanding system for private-by-design voice interfaces

Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, Maël Primet, Joseph Dureau

1805.10190

Deep backward schemes for high-dimensional nonlinear PDEs

Côme Huré, Huyên Pham, Xavier Warin

1902.01599

NeuroNER: an easy-to-use program for named-entity recognition based on neural networks

Franck Dernoncourt, Ji Young Lee, Peter Szolovits

1705.05487

SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention

Róbert Csordás, Piotr Piękos, Kazuki Irie, Jürgen Schmidhuber

2312.07987

Multimodal Convolutional Neural Networks for Matching Image and Sentence

Lin Ma, Zhengdong Lu, Lifeng Shang, Hang Li

1504.06063

Neural networks-based backward scheme for fully nonlinear PDEs

Huyen Pham, Xavier Warin, Maximilien Germain

1908.00412

Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning

Felipe Petroski Such, Vashisht Madhavan, Edoardo Conti, Joel Lehman, Kenneth O. Stanley, Jeff Clune

1712.06567

A Fine-Grained Spectral Perspective on Neural Networks

Greg Yang, Hadi Salman

1907.10599

Sparse Networks from Scratch: Faster Training without Losing Performance

Tim Dettmers, Luke Zettlemoyer

1907.04840

Neural Photo Editing with Introspective Adversarial Networks

Andrew Brock, Theodore Lim, J. M. Ritchie, Nick Weston

1609.07093

Neural Code Comprehension: A Learnable Representation of Code Semantics

Tal Ben-Nun, Alice Shoshana Jakobovits, Torsten Hoefler

1806.07336

Capacity and Trainability in Recurrent Neural Networks

Jasmine Collins, Jascha Sohl-Dickstein, David Sussillo

1611.09913

ACDC: A Structured Efficient Linear Layer

Marcin Moczulski, Misha Denil, Jeremy Appleyard, Nando de Freitas

1511.05946

Accelerating Very Deep Convolutional Networks for Classification and Detection

Xiangyu Zhang, Jianhua Zou, Kaiming He, Jian Sun

1505.06798

Graphite: Iterative Generative Modeling of Graphs

Aditya Grover, Aaron Zweig, Stefano Ermon

1803.10459

The Variational Gaussian Process

Dustin Tran, Rajesh Ranganath, David M. Blei

1511.06499

The power of deeper networks for expressing natural functions

David Rolnick, Max Tegmark

1705.05502

Task Agnostic Continual Learning via Meta Learning

Xu He, Jakub Sygnowski, Alexandre Galashov, Andrei A. Rusu, Yee Whye Teh, Razvan Pascanu

1906.05201

Rademacher Complexity for Adversarially Robust Generalization

Dong Yin, Kannan Ramchandran, Peter Bartlett

1810.11914

Large-scale Point Cloud Semantic Segmentation with Superpoint Graphs

Loic Landrieu, Martin Simonovsky

1711.09869

A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks

Dan Hendrycks, Kevin Gimpel

1610.02136

Lookahead Optimizer: k steps forward, 1 step back

Michael R. Zhang, James Lucas, Geoffrey Hinton, Jimmy Ba

1907.08610

Deep API Learning

Xiaodong Gu, Hongyu Zhang, Dongmei Zhang, Sunghun Kim

1605.08535

Overlearning Reveals Sensitive Attributes

Congzheng Song, Vitaly Shmatikov

1905.11742

De-identification of Patient Notes with Recurrent Neural Networks

Franck Dernoncourt, Ji Young Lee, Ozlem Uzuner, Peter Szolovits

1606.03475

Neural Lyapunov Control

Ya-Chien Chang, Nima Roohi, Sicun Gao

2005.00611

Transformer-based Online CTC/attention End-to-End Speech Recognition Architecture

Haoran Miao, Gaofeng Cheng, Changfeng Gao, Pengyuan Zhang, Yonghong Yan

2001.08290

Flexible Neural Representation for Physics Prediction

Damian Mrowca, Chengxu Zhuang, Elias Wang, Nick Haber, Li Fei-Fei, Joshua B. Tenenbaum, Daniel L. K. Yamins

1806.08047

Choose Your Programming Copilot: A Comparison of the Program Synthesis Performance of GitHub Copilot and Genetic Programming

Dominik Sobania, Martin Briesch, Franz Rothlauf

2111.07875

Model compression via distillation and quantization

Antonio Polino, Razvan Pascanu, Dan Alistarh

1802.05668

Frustratingly Short Attention Spans in Neural Language Modeling

Michał Daniluk, Tim Rocktäschel, Johannes Welbl, Sebastian Riedel

1702.04521

A Hierarchical Recurrent Encoder-Decoder For Generative Context-Aware Query Suggestion

Alessandro Sordoni, Yoshua Bengio, Hossein Vahabi, Christina Lioma, Jakob G. Simonsen, Jian-Yun Nie

1507.02221

Synthesizing the preferred inputs for neurons in neural networks via deep generator networks

Anh Nguyen, Alexey Dosovitskiy, Jason Yosinski, Thomas Brox, Jeff Clune

1605.09304

pix2code: Generating Code from a Graphical User Interface Screenshot

Tony Beltramelli

1705.07962

Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning Models

Wieland Brendel, Jonas Rauber, Matthias Bethge

1712.04248

One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers

Ari S. Morcos, Haonan Yu, Michela Paganini, Yuandong Tian

1906.02773

Multi-task Learning of Pairwise Sequence Classification Tasks Over Disparate Label Spaces

Isabelle Augenstein, Sebastian Ruder, Anders Søgaard

1802.09913

Entity Abstraction in Visual Model-Based Reinforcement Learning

Rishi Veerapaneni, John D. Co-Reyes, Michael Chang, Michael Janner, Chelsea Finn, Jiajun Wu, Joshua B. Tenenbaum, Sergey Levine

1910.12827

Multiresolution Recurrent Neural Networks: An Application to Dialogue Response Generation

Iulian Vlad Serban, Tim Klinger, Gerald Tesauro, Kartik Talamadupula, Bowen Zhou, Yoshua Bengio, Aaron Courville

1606.00776

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution

Chrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero, Tim Rocktäschel

2309.16797

A Closer Look at Deep Policy Gradients

Andrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, Aleksander Madry

1811.02553

Random feedback weights support learning in deep neural networks

Timothy P. Lillicrap, Daniel Cownden, Douglas B. Tweed, Colin J. Akerman

1411.0247

A Simple Neural Attentive Meta-Learner

Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, Pieter Abbeel

1707.03141

Tensor Programs IVb: Adaptive Optimization in the Infinite-Width Limit

Greg Yang, Etai Littwin

2308.01814

A Survey on Methods and Theories of Quantized Neural Networks

Yunhui Guo

1808.04752

Neural Machine Translation with Recurrent Attention Modeling

Zichao Yang, Zhiting Hu, Yuntian Deng, Chris Dyer, Alex Smola

1607.05108

Explaining Predictions of Non-Linear Classifiers in NLP

Leila Arras, Franziska Horn, Grégoire Montavon, Klaus-Robert Müller, Wojciech Samek

1606.07298

Provably Optimal Memory Capacity for Modern Hopfield Models: Transformer-Compatible Dense Associative Memories as Spherical Codes

Jerry Yao-Chieh Hu, Dennis Wu, Han Liu

2410.23126

Reverse Curriculum Generation for Reinforcement Learning

Carlos Florensa, David Held, Markus Wulfmeier, Michael Zhang, Pieter Abbeel

1707.05300

Curiosity Driven Exploration of Learned Disentangled Goal Spaces

Adrien Laversanne-Finot, Alexandre Péré, Pierre-Yves Oudeyer

1807.01521

LLMatic: Neural Architecture Search via Large Language Models and Quality Diversity Optimization

Muhammad U. Nasir, Sam Earle, Christopher Cleghorn, Steven James, Julian Togelius

2306.01102

How Transferable are Neural Networks in NLP Applications?

Lili Mou, Zhao Meng, Rui Yan, Ge Li, Yan Xu, Lu Zhang, Zhi Jin

1603.06111

Imposing higher-level Structure in Polyphonic Music Generation using Convolutional Restricted Boltzmann Machines and Constraints

Stefan Lattner, Maarten Grachten, Gerhard Widmer

1612.04742

Correlational Neural Networks

Sarath Chandar, Mitesh M. Khapra, Hugo Larochelle, Balaraman Ravindran

1504.07225

Dynamic Edge-Conditioned Filters in Convolutional Neural Networks on Graphs

Martin Simonovsky, Nikos Komodakis

1704.02901

SEGAN: Speech Enhancement Generative Adversarial Network

Santiago Pascual, Antonio Bonafonte, Joan Serrà

1703.09452

Convolutional Recurrent Neural Networks for Music Classification

Keunwoo Choi, George Fazekas, Mark Sandler, Kyunghyun Cho

1609.04243

Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions

Rui Wang, Joel Lehman, Jeff Clune, Kenneth O. Stanley

1901.01753

SE3-Pose-Nets: Structured Deep Dynamics Models for Visuomotor Planning and Control

Arunkumar Byravan, Felix Leeb, Franziska Meier, Dieter Fox

1710.00489

Self-Adaptive Hierarchical Sentence Model

Han Zhao, Zhengdong Lu, Pascal Poupart

1504.05070

Learning Visual Question Answering by Bootstrapping Hard Attention

Mateusz Malinowski, Carl Doersch, Adam Santoro, Peter Battaglia

1808.00300

Failures of Gradient-Based Deep Learning

Shai Shalev-Shwartz, Ohad Shamir, Shaked Shammah

1703.07950

EvoPrompting: Language Models for Code-Level Neural Architecture Search

Angelica Chen, David M. Dohan, David R. So

2302.14838

Semantic Parsing with Semi-Supervised Sequential Autoencoders

Tomáš Kočiský, Gábor Melis, Edward Grefenstette, Chris Dyer, Wang Ling, Phil Blunsom, Karl Moritz Hermann

1609.09315

On the Expressive Power of Deep Polynomial Neural Networks

Joe Kileel, Matthew Trager, Joan Bruna

1905.12207

3D Point Capsule Networks

Yongheng Zhao, Tolga Birdal, Haowen Deng, Federico Tombari

1812.10775

A Systematic DNN Weight Pruning Framework using Alternating Direction Method of Multipliers

Tianyun Zhang, Shaokai Ye, Kaiqi Zhang, Jian Tang, Wujie Wen, Makan Fardad, Yanzhi Wang

1804.03294

Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures

Hengyuan Hu, Rui Peng, Yu-Wing Tai, Chi-Keung Tang

1607.03250

Multilingual Image Description with Neural Sequence Models

Desmond Elliott, Stella Frank, Eva Hasler

1510.04709

Training Deep Neural Networks on Noisy Labels with Bootstrapping

Scott Reed, Honglak Lee, Dragomir Anguelov, Christian Szegedy, Dumitru Erhan, Andrew Rabinovich

1412.6596

Continual Pre-training of Language Models

Zixuan Ke, Yijia Shao, Haowei Lin, Tatsuya Konishi, Gyuhak Kim, Bing Liu

2302.03241

Using Trusted Data to Train Deep Networks on Labels Corrupted by Severe Noise

Dan Hendrycks, Mantas Mazeika, Duncan Wilson, Kevin Gimpel

1802.05300

Twin Regularization for online speech recognition

Mirco Ravanelli, Dmitriy Serdyuk, Yoshua Bengio

1804.05374

Zero-shot Sequence Labeling: Transferring Knowledge from Sentences to Tokens

Marek Rei, Anders Søgaard

1805.02214

Resource-Efficient Neural Architect

Yanqi Zhou, Siavash Ebrahimi, Sercan Ö. Arık, Haonan Yu, Hairong Liu, Greg Diamos

1806.07912

Grow and Prune Compact, Fast, and Accurate LSTMs

Xiaoliang Dai, Hongxu Yin, Niraj K. Jha

1805.11797

A New 2.5D Representation for Lymph Node Detection using Random Sets of Deep Convolutional Neural Network Observations

Holger R. Roth, Le Lu, Ari Seff, Kevin M. Cherry, Joanne Hoffman, Shijun Wang, Jiamin Liu, Evrim Turkbey, Ronald M. Summers

1406.2639

Finer Grained Entity Typing with TypeNet

Shikhar Murty, Patrick Verga, Luke Vilnis, Andrew McCallum

1711.05795

A Latent Variable Recurrent Neural Network for Discourse Relation Language Models

Yangfeng Ji, Gholamreza Haffari, Jacob Eisenstein

1603.01913

Inference with Deep Generative Priors in High Dimensions

Parthe Pandit, Mojtaba Sahraee-Ardakan, Sundeep Rangan, Philip Schniter, Alyson K. Fletcher

1911.03409

Reconstructing Training Data from Trained Neural Networks

Niv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir, Michal Irani

2206.07758

Learning from LDA using Deep Neural Networks

Dongxu Zhang, Tianyi Luo, Dong Wang, Rong Liu

1508.01011

BranchyNet: Fast Inference via Early Exiting from Deep Neural Networks

Surat Teerapittayanon, Bradley McDanel, H. T. Kung

1709.01686

Improvements to deep convolutional neural networks for LVCSR

Tara N. Sainath, Brian Kingsbury, Abdel-rahman Mohamed, George E. Dahl, George Saon, Hagen Soltau, Tomas Beran, Aleksandr Y. Aravkin, Bhuvana Ramabhadran

1309.1501

BlackOut: Speeding up Recurrent Neural Network Language Models With Very Large Vocabularies

Shihao Ji, S. V. N. Vishwanathan, Nadathur Satish, Michael J. Anderson, Pradeep Dubey

1511.06909

When Face Recognition Meets with Deep Learning: an Evaluation of Convolutional Neural Networks for Face Recognition

Guosheng Hu, Yongxin Yang, Dong Yi, Josef Kittler, William Christmas, Stan Z. Li, Timothy Hospedales

1504.02351

Understanding Generalization through Visualizations

W. Ronny Huang, Zeyad Emam, Micah Goldblum, Liam Fowl, J. K. Terry, Furong Huang, Tom Goldstein

1906.03291

Low-Dose CT with a Residual Encoder-Decoder Convolutional Neural Network (RED-CNN)

Hu Chen, Yi Zhang, Mannudeep K. Kalra, Feng Lin, Yang Chen, Peixi Liao, Jiliu Zhou, Ge Wang

1702.00288

Smooth Neighbors on Teacher Graphs for Semi-supervised Learning

Yucen Luo, Jun Zhu, Mengxi Li, Yong Ren, Bo Zhang

1711.00258

Self-supervised learning through the eyes of a child

A. Emin Orhan, Vaibhav V. Gupta, Brenden M. Lake

2007.16189

Compression of Neural Machine Translation Models via Pruning

Abigail See, Minh-Thang Luong, Christopher D. Manning

1606.09274

Planning to Explore via Self-Supervised World Models

Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, Deepak Pathak

2005.05960

The asymptotic spectrum of the Hessian of DNN throughout training

Arthur Jacot, Franck Gabriel, Clément Hongler

1910.02875

Consensus Attention-based Neural Networks for Chinese Reading Comprehension

Yiming Cui, Ting Liu, Zhipeng Chen, Shijin Wang, Guoping Hu

1607.02250

Scalable agent alignment via reward modeling: a research direction

Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, Shane Legg

1811.07871

Parameter Hub: a Rack-Scale Parameter Server for Distributed Deep Neural Network Training

Liang Luo, Jacob Nelson, Luis Ceze, Amar Phanishayee, Arvind Krishnamurthy

1805.07891

Randomness In Neural Network Training: Characterizing The Impact of Tooling

Donglin Zhuang, Xingyao Zhang, Shuaiwen Leon Song, Sara Hooker

2106.11872

DeepAM: Migrate APIs with Multi-modal Sequence to Sequence Learning

Xiaodong Gu, Hongyu Zhang, Dongmei Zhang, Sunghun Kim

1704.07734

Automatic Rule Extraction from Long Short Term Memory Networks

W. James Murdoch, Arthur Szlam

1702.02540

Transferring Knowledge from a RNN to a DNN

William Chan, Nan Rosemary Ke, Ian Lane

1504.01483

Accelerating Deep Convolutional Networks using low-precision and sparsity

Ganesh Venkatesh, Eriko Nurvitadhi, Debbie Marr

1610.00324

Hello Edge: Keyword Spotting on Microcontrollers

Yundong Zhang, Naveen Suda, Liangzhen Lai, Vikas Chandra

1711.07128

ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution Blocks

Xiaohan Ding, Yuchen Guo, Guiguang Ding, Jungong Han

1908.03930

Building DNN Acoustic Models for Large Vocabulary Speech Recognition

Andrew L. Maas, Peng Qi, Ziang Xie, Awni Y. Hannun, Christopher T. Lengerich, Daniel Jurafsky, Andrew Y. Ng

1406.7806

SqueezeNext: Hardware-Aware Neural Network Design

Amir Gholami, Kiseok Kwon, Bichen Wu, Zizheng Tai, Xiangyu Yue, Peter Jin, Sicheng Zhao, Kurt Keutzer

1803.10615

RETURNN: The RWTH Extensible Training framework for Universal Recurrent Neural Networks

Patrick Doetsch, Albert Zeyer, Paul Voigtlaender, Ilya Kulikov, Ralf Schlüter, Hermann Ney

1608.00895

Deep Reinforcement Learning in Parameterized Action Space

Matthew Hausknecht, Peter Stone

1511.04143

Context-Free Transductions with Neural Stacks

Yiding Hao, William Merrill, Dana Angluin, Robert Frank, Noah Amsel, Andrew Benz, Simon Mendelsohn

1809.02836

Powerset multi-class cross entropy loss for neural speaker diarization

Alexis Plaquet, Hervé Bredin

2310.13025

PSB2: The Second Program Synthesis Benchmark Suite

Thomas Helmuth, Peter Kelly

2106.06086

Deep learning generalizes because the parameter-function map is biased towards simple functions

Guillermo Valle-Pérez, Chico Q. Camargo, Ard A. Louis

1805.08522

Evolution through Large Models

Joel Lehman, Jonathan Gordon, Shawn Jain, Kamal Ndousse, Cathy Yeh, Kenneth O. Stanley

2206.08896

Machine Learning on Graphs: A Model and Comprehensive Taxonomy

Ines Chami, Sami Abu-El-Haija, Bryan Perozzi, Christopher Ré, Kevin Murphy

2005.03675

Nonparametric Modern Hopfield Models

Jerry Yao-Chieh Hu, Bo-Yu Chen, Dennis Wu, Feng Ruan, Han Liu

2404.03900

Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNs

Jonathan Frankle, David J. Schwab, Ari S. Morcos

2003.00152

Deep Speech: Scaling up end-to-end speech recognition

Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, Andrew Y. Ng

1412.5567

Deep Learning Approximation for Stochastic Control Problems

Jiequn Han, Weinan E

1611.07422

Ansor: Generating High-Performance Tensor Programs for Deep Learning

Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, Ion Stoica

2006.06762

Deep Unsupervised Clustering with Gaussian Mixture Variational Autoencoders

Nat Dilokthanakul, Pedro A. M. Mediano, Marta Garnelo, Matthew C. H. Lee, Hugh Salimbeni, Kai Arulkumaran, Murray Shanahan

1611.02648

Memory-Augmented Recurrent Neural Networks Can Learn Generalized Dyck Languages

Mirac Suzgun, Sebastian Gehrmann, Yonatan Belinkov, Stuart M. Shieber

1911.03329

Numeracy for Language Models: Evaluating and Improving their Ability to Predict Numbers

Georgios P. Spithourakis, Sebastian Riedel

1805.08154

Learning to Combine Top-Down and Bottom-Up Signals in Recurrent Neural Networks with Attention over Modules

Sarthak Mittal, Alex Lamb, Anirudh Goyal, Vikram Voleti, Murray Shanahan, Guillaume Lajoie, Michael Mozer, Yoshua Bengio

2006.16981

Modular Multitask Reinforcement Learning with Policy Sketches

Jacob Andreas, Dan Klein, Sergey Levine

1611.01796

Segmental Recurrent Neural Networks for End-to-end Speech Recognition

Liang Lu, Lingpeng Kong, Chris Dyer, Noah A. Smith, Steve Renals

1603.00223

Associative Long Short-Term Memory

Ivo Danihelka, Greg Wayne, Benigno Uria, Nal Kalchbrenner, Alex Graves

1602.03032

An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks

Ian J. Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, Yoshua Bengio

1312.6211

Byzantine-Tolerant Machine Learning

Peva Blanchard, El Mahdi El Mhamdi, Rachid Guerraoui, Julien Stainer

1703.02757

Hopfield Networks is All You Need

Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlović, Geir Kjetil Sandve, Victor Greiff, David Kreil, Michael Kopp, Günter Klambauer, Johannes Brandstetter, Sepp Hochreiter

2008.02217

Convergence of gradient descent for deep neural networks

Sourav Chatterjee

2203.16462

Weisfeiler and Leman go Machine Learning: The Story so far

Christopher Morris, Yaron Lipman, Haggai Maron, Bastian Rieck, Nils M. Kriege, Martin Grohe, Matthias Fey, Karsten Borgwardt

2112.09992

Generating Focussed Molecule Libraries for Drug Discovery with Recurrent Neural Networks

Marwin H. S. Segler, Thierry Kogej, Christian Tyrchan, Mark P. Waller

1701.01329

Understanding Anomaly Detection with Deep Invertible Networks through Hierarchies of Distributions and Features

Robin Tibor Schirrmeister, Yuxuan Zhou, Tonio Ball, Dan Zhang

2006.10848

Efficient Sparse-Winograd Convolutional Neural Networks

Xingyu Liu, Jeff Pool, Song Han, William J. Dally

1802.06367

Data-Free Network Quantization With Adversarial Knowledge Distillation

Yoojin Choi, Jihwan Choi, Mostafa El-Khamy, Jungwon Lee

2005.04136

The Power of the Weisfeiler-Leman Algorithm for Machine Learning with Graphs

Christopher Morris, Matthias Fey, Nils M. Kriege

2105.05911

A Unified Framework of Online Learning Algorithms for Training Recurrent Neural Networks

Owen Marschall, Kyunghyun Cho, Cristina Savin

1907.02649

Mixed Precision Training of Convolutional Neural Networks using Integer Operations

Dipankar Das, Naveen Mellempudi, Dheevatsa Mudigere, Dhiraj Kalamkar, Sasikanth Avancha, Kunal Banerjee, Srinivas Sridharan, Karthik Vaidyanathan, Bharat Kaul, Evangelos Georganas, Alexander Heinecke, Pradeep Dubey, Jesus Corbal, Nikita Shustrov, Roma Dubtsov, Evarist Fomenko, Vadim Pirogov

1802.00930

Session-based Recommendations with Recurrent Neural Networks

Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, Domonkos Tikk

1511.06939

Equivariance Through Parameter-Sharing

Siamak Ravanbakhsh, Jeff Schneider, Barnabas Poczos

1702.08389

PREMA: A Predictive Multi-task Scheduling Algorithm For Preemptible Neural Processing Units

Yujeong Choi, Minsoo Rhu

1909.04548

PDE-GCN: Novel Architectures for Graph Neural Networks Motivated by Partial Differential Equations

Moshe Eliasof, Eldad Haber, Eran Treister

2108.01938

Attention Strategies for Multi-Source Sequence-to-Sequence Learning

Jindřich Libovický, Jindřich Helcl

1704.06567

Fast-Slow Recurrent Neural Networks

Asier Mujika, Florian Meier, Angelika Steger

1705.08639

Tradeoffs between Convergence Speed and Reconstruction Accuracy in Inverse Problems

Raja Giryes, Yonina C. Eldar, Alex M. Bronstein, Guillermo Sapiro

1605.09232

Elastic Graph Neural Networks

Xiaorui Liu, Wei Jin, Yao Ma, Yaxin Li, Hua Liu, Yiqi Wang, Ming Yan, Jiliang Tang

2107.06996

TasNet: time-domain audio separation network for real-time, single-channel speech separation

Yi Luo, Nima Mesgarani

1711.00541

Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches

Maurizio Ferrari Dacrema, Paolo Cremonesi, Dietmar Jannach

1907.06902

Towards Deep Learning Models Resistant to Adversarial Attacks

Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, Adrian Vladu

1706.06083

Multi-turn Dialogue Response Generation in an Adversarial Learning Framework

Oluwatobi Olabiyi, Alan Salimov, Anish Khazane, Erik T. Mueller

1805.11752

Language Model Crossover: Variation through Few-Shot Prompting

Elliot Meyerson, Mark J. Nelson, Herbie Bradley, Adam Gaier, Arash Moradi, Amy K. Hoover, Joel Lehman

2302.12170

The NetHack Learning Environment

Heinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, Tim Rocktäschel

2006.13760

Adapting a Language Model While Preserving its General Knowledge

Zixuan Ke, Yijia Shao, Haowei Lin, Hu Xu, Lei Shu, Bing Liu

2301.08986

Super-Linear Gate and Super-Quadratic Wire Lower Bounds for Depth-Two and Depth-Three Threshold Circuits

Daniel M. Kane, Ryan Williams

1511.07860

CMSIS-NN: Efficient Neural Network Kernels for Arm Cortex-M CPUs

Liangzhen Lai, Naveen Suda, Vikas Chandra

1801.06601

Listen, Attend, and Walk: Neural Mapping of Navigational Instructions to Action Sequences

Hongyuan Mei, Mohit Bansal, Matthew R. Walter

1506.04089

Discourse-Based Objectives for Fast Unsupervised Sentence Representation Learning

Yacine Jernite, Samuel R. Bowman, David Sontag

1705.00557

End-to-End ASR-free Keyword Search from Speech

Kartik Audhkhasi, Andrew Rosenberg, Abhinav Sethy, Bhuvana Ramabhadran, Brian Kingsbury

1701.04313

A Large Contextual Dataset for Classification, Detection and Counting of Cars with Deep Learning

T. Nathan Mundhenk, Goran Konjevod, Wesam A. Sakla, Kofi Boakye

1609.04453

Overcoming catastrophic forgetting with hard attention to the task

Joan Serrà, Dídac Surís, Marius Miron, Alexandros Karatzoglou

1801.01423

Domain Adaptation via Teacher-Student Learning for End-to-End Speech Recognition

Zhong Meng, Jinyu Li, Yashesh Gaur, Yifan Gong

2001.01798

PoCoNet: Better Speech Enhancement with Frequency-Positional Embeddings, Semi-Supervised Conversational Data, and Biased Loss

Umut Isik, Ritwik Giri, Neerad Phansalkar, Jean-Marc Valin, Karim Helwani, Arvindh Krishnaswamy

2008.04470

Transforming Exploratory Creativity with DeLeNoX

Antonios Liapis, Hector P. Martinez, Julian Togelius, Georgios N. Yannakakis

2103.11715

Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification

Justin Salamon, Juan Pablo Bello

1608.04363

Global convergence of ResNets: From finite to infinite width using linear parameterization

Raphaël Barboni, Gabriel Peyré, François-Xavier Vialard

2112.05531

Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments

Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, Igor Mordatch

1706.02275

The Mythos of Model Interpretability

Zachary C. Lipton

1606.03490

Understanding Synthetic Gradients and Decoupled Neural Interfaces

Wojciech Marian Czarnecki, Grzegorz Świrszcz, Max Jaderberg, Simon Osindero, Oriol Vinyals, Koray Kavukcuoglu

1703.00522

What Matters for Adversarial Imitation Learning?

Manu Orsini, Anton Raichuk, Léonard Hussenot, Damien Vincent, Robert Dadashi, Sertan Girgin, Matthieu Geist, Olivier Bachem, Olivier Pietquin, Marcin Andrychowicz

2106.00672

Goal-conditioned Imitation Learning

Yiming Ding, Carlos Florensa, Mariano Phielipp, Pieter Abbeel

1906.05838

A Practical Sparse Approximation for Real Time Recurrent Learning

Jacob Menick, Erich Elsen, Utku Evci, Simon Osindero, Karen Simonyan, Alex Graves

2006.07232

A Convergence Theory for Deep Learning via Over-Parameterization

Zeyuan Allen-Zhu, Yuanzhi Li, Zhao Song

1811.03962

How to Construct Deep Recurrent Neural Networks

Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Yoshua Bengio

1312.6026

Long Short-Term Memory Based Recurrent Neural Network Architectures for Large Vocabulary Speech Recognition

Haşim Sak, Andrew Senior, Françoise Beaufays

1402.1128

Learned-Norm Pooling for Deep Feedforward and Recurrent Neural Networks

Caglar Gulcehre, Kyunghyun Cho, Razvan Pascanu, Yoshua Bengio

1311.1780

On Fast Dropout and its Applicability to Recurrent Networks

Justin Bayer, Christian Osendorfer, Daniela Korhammer, Nutan Chen, Sebastian Urban, Patrick van der Smagt

1311.0701

Metric-Free Natural Gradient for Joint-Training of Boltzmann Machines

Guillaume Desjardins, Razvan Pascanu, Aaron Courville, Yoshua Bengio

1301.3545

Knowledge Matters: Importance of Prior Information for Optimization

Çağlar Gülçehre, Yoshua Bengio

1301.4083

Training Neural Networks with Stochastic Hessian-Free Optimization

Ryan Kiros

1301.3641

Batch Normalized Recurrent Neural Networks

César Laurent, Gabriel Pereyra, Philémon Brakel, Ying Zhang, Yoshua Bengio

1510.01378

Programming with a Differentiable Forth Interpreter

Matko Bošnjak, Tim Rocktäschel, Jason Naradowsky, Sebastian Riedel

1605.06640

Adversarial Transformation Networks: Learning to Generate Adversarial Examples

Shumeet Baluja, Ian Fischer

1703.09387

Unmasking Clever Hans Predictors and Assessing What Machines Really Learn

Sebastian Lapuschkin, Stephan Wäldchen, Alexander Binder, Grégoire Montavon, Wojciech Samek, Klaus-Robert Müller

1902.10178

Neural Speech Recognizer: Acoustic-to-Word LSTM Model for Large Vocabulary Speech Recognition

Hagen Soltau, Hank Liao, Hasim Sak

1610.09975

Building competitive direct acoustics-to-word models for English conversational speech recognition

Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Michael Picheny

1712.03133

Exploring Neural Transducers for End-to-End Speech Recognition

Eric Battenberg, Jitong Chen, Rewon Child, Adam Coates, Yashesh Gaur, Yi Li, Hairong Liu, Sanjeev Satheesh, David Seetapun, Anuroop Sriram, Zhenyao Zhu

1707.07413

Population Based Training of Neural Networks

Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M. Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, Chrisantha Fernando, Koray Kavukcuoglu

1711.09846

Deep Learning with Limited Numerical Precision

Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, Pritish Narayanan

1502.02551

Structural Attention Neural Networks for improved sentiment analysis

Filippos Kokkinos, Alexandros Potamianos

1701.01811

Ternary Neural Networks for Resource-Efficient AI Applications

Hande Alemdar, Vincent Leroy, Adrien Prost-Boucle, Frédéric Pétrot

1609.00222

The Lottery Ticket Hypothesis for Pre-trained BERT Networks

Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang Wang, Michael Carbin

2007.12223

Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results

Antti Tarvainen, Harri Valpola

1703.01780

Difference Target Propagation

Dong-Hyun Lee, Saizheng Zhang, Asja Fischer, Yoshua Bengio

1412.7525

Evolutionary Population Curriculum for Scaling Multi-Agent Reinforcement Learning

Qian Long, Zihan Zhou, Abhibav Gupta, Fei Fang, Yi Wu, Xiaolong Wang

2003.10423

Exponential Convergence Time of Gradient Descent for One-Dimensional Deep Linear Neural Networks

Ohad Shamir

1809.08587

Latent Multi-task Architecture Learning

Sebastian Ruder, Joachim Bingel, Isabelle Augenstein, Anders Søgaard

1705.08142