Fair DARTS: Eliminating Unfair Advantages in Differentiable Architecture Search

Xiangxiang Chu, Tianbao Zhou, Bo Zhang, Jixiang Li

Introduction

In the wake of the DARTS’s open-sourcing , a diverse number of its variants emerge in the neural architecture search community. Some of them extend its use in higher-level architecture search spaces with performance awareness in mind , some learn a stochastic distribution instead of architectural parameters , and others offer remedies on discovering its lack of robustness .

In spite of these endeavors, the aggregation of skip connections in DARTS that noticed by has not been solved with perfection. Observing that the aggregation leads to a dramatic performance collapse for the resulting architecture, P-DARTS utilizes dropout as a workaround to restrict the number of skip connections during optimization. DARTS+ directly puts a hard limit of two skip-connections per cell. RobustDARTS finds out that these solutions coincide with high validation loss curvatures. To some extent, these approaches consider the poor-performing models as impurities from the solution set, for which they intervene in the training process to filter them out.

On the contrary, we extend the solution set and revise the optimization process so that aggregation of skip connections no longer causes the collapse. Moreover, there remains a discrepancy problem when discretizing continuous architecture encodings. DARTS leaves it as future work, but till now it has not been deeply studied. We reiterate the basic premise of DARTS is that the continuous solution approximates a one-hot encoding. Intuitively, the smaller discrepancies are, the more consistent it will be when we transform a continuous solution back to a discrete one. We summarize our contributions as follows:

Firstly, we disclose the root cause that leads to the collapse of DARTS, which we later define as an unfair advantage that drives skip connections into a monopoly state in exclusive competition. These two indispensable factors work together to induce a performance collapse. Moreover, if either of the two conditions is broken, the collapse disappears.

Secondly, we propose the first collaborative competition approach by offering each operation an independent architectural weight. The unfair advantage no longer prevails as we break the second factor. Furthermore, to address the discrepancy between the continuous architecture encoding and the derived discrete one in our method, we propose a novel auxiliary loss, called zero-one loss, to steer architectural weights towards their extremities, that is, either completely enabled or disabled. The discrepancy thus decreases to its minimum.

Thirdly, based on the root cause of the collapse, we provide a unified perspective to view current DARTS cures for skip connections’ aggregation. The majority of these works either make use of dropout on skip connections , or play with the later termed boundary epoch by different early-stopping strategies . They can all be regarded as preventing the first factor from taking effect. Moreover, as a direct application, we can derive a hypothesis that adding Gaussian noise also disrupts the unfairness, which is later proved to be effective.

Lastly, we conduct thorough experiments in two widely used search spaces in both proxy and proxyless ways. Results show that our method can escape from performance collapse. We also achieve state-of-the-art networks on CIFAR-10 and ImageNet.

Related Work

Lately, neural architecture search has grown as a well-formed methodology to discover networks for various deep learning tasks. Endeavors have been made to reduce the enormous searching overhead with the weight-sharing mechanism . Especially in DARTS , a nested gradient-descent algorithm is exploited to search for the graphical representation of architectures, which is born from gradient-based hyperparameter optimization .

Due to the limit of the DARTS search space, ProxylessNAS and FBNet apply DARTS in much larger search spaces based on MobileNetV2 . ProxylessNAS also differs from DARTS in its supernet training process, where only two paths are activated, based on the assumption that one path is the best amongst all should be better than any single one. From a fairness point of view, as only two paths enhance their ability (get parameters updated) while others remain unchanged, it implicitly creates a bias. FBNet , SNAS and GDAS utilize the differentiable Gumbel Softmax to mimic one-hot encoding. However, the one-hot nature implies an exclusive competition, which risks being exploited by unfair advantages.

Superficially, the most relevant work to ours is RobustDARTS . Under several simplified search spaces, they state that the found solutions generalize poorly when they coincide with high validation loss curvature, where the supernet with an excessive number of skip connections happens to be such a solution. Based on this observation, they impose early-stop regularization by tracking the largest eigenvalue. Instead, our method doesn’t need to perform early stopping.

The Downside of DARTS

In this section, we aim to excavate the disadvantages of DARTS that possibly impede the searching performance. We first prepare a minimum background.

For the case of convolutional neural networks, DARTS searches for a normal cell and a reduction cell to build up the final architecture. A cell is represented as a directed acyclic graph (DAG) of NN nodes in sequential order. Each node stands for a feature map. The edge ei,je_{i,j} from node ii to jj operates on the input feature xix_{i} and its output is denoted as oi,j(xi)o_{i,j}(x_{i}). The intermediate node jj gathers all inputs from the incoming edges,

Let O={oi,j1,oi,j2,...,oi,jM}\mathcal{O}=\{o_{i,j}^{1},o_{i,j}^{2},...,o_{i,j}^{M}\} be the set of MM candidate operations on edge ei,je_{i,j}. DARTS relaxes this categorical choice to a softmax over all operations in O\mathcal{O} to form a mixed output:

where each operation oi,jo_{i,j} is associated with a continuous coefficient αoi,j\alpha_{o_{i,j}}. Regarding edge ei,je_{i,j}, this softmax is utilized to approximate one-hot encoding βi,j=(βoi,j1,βoi,j2,...,βoi,jM)\beta_{i,j}=(\beta_{o_{i,j}^{1}},\beta_{o_{i,j}^{2}},...,\beta_{o_{i,j}^{M}}). Formally, let αoi,j\alpha_{o_{i,j}} denote the architectural weights vector (αoi,j1\alpha_{o_{i,j}^{1}}, αoi,j2\alpha_{o_{i,j}^{2}}, …, αoi,jM\alpha_{o_{i,j}^{M}}). DARTS thus assumes the following as a valid approximation,

The architecture search problem is reduced to learning α∗\alpha^{*} and network weights w∗w^{*} that minimize the validation loss Lval(w∗,α∗)\mathcal{L}_{val}(w^{*},\alpha^{*}). DARTS resolves this problem with a bi-level optimization,

We also adopt two common search spaces, the DARTS search space (S1S_{1}) and the ProxylessNAS search space (S2S_{2}) with minor modifications. More details are given in Section 9 (supplementary).

In S2S_{2}, the output of the ll-th layer is a softmax-weighted summation of NN choices. Formally, it can be written as

2 Performance Collapse Caused by Intractable Skip Connections

DARTS suffers from significant performance decay when skip connections become dominant . It was described as a competition-and-cooperation issue in the bi-level optimization . Still, the reason behind this behavior is not clear, we hereby provide a different perspective.

First, to confirm this issue, we run DARTS k=4k=4 times with different random seeds. Following DARTS, we select 8 top-performing operations per cell (2 each for 4 intermediate nodes). Here we say one operation is dominant if it has top-2 softmax(α)softmax(\alpha) among all incoming edges’ candidates of a certain node. The results are shown in Fig. 2. In the beginning, all operations are given the same opportunity. As the over-parameterized network gradually converges, there is an evident aggregation of skip connections after 20 epochs (5 out of 8 in an extreme case).

When we utilize DARTS directly on ImageNet in S2S_{2}, which is a single branch architecture, the same phenomenon rigorously reappears. The number of dominant skip-connections (highest softmax(α)softmax(\alpha) among all operations in that layer) steadily increases and reaches 11 out of 19 layers in the end, which is shown on the left of Fig. 3.

But why is this happening? The underlying reasons are rarely discussed in depth. A brief and superficial analysis regarding information flow is given in . However, we claim that the reason for excessive skip connections is from exclusive competition among various operations. In Equation 2 and Equation 5, the skip connection is softmax-weighted and added to the output, which resembles a basic residual module as in ResNet . While this module greatly benefits the training, the architectural weight of a skip connection increases much faster than its competitors. Moreover, the softmax operation inherently provides an exclusive competition since increasing one is at the cost of suppressing others. As a result, skip connections become gradually dominant during optimization. We have to keep in mind that skip connection works well because it is in cooperation with convolutions . However, DARTS picks the top-performing one (skip connection here) and discards its collaborator (convolution), which results in a degenerate model.

We further study this effect from the experiments on CIFAR-10 by recording the competition progress in Fig. 3. The derived model has 8 skip connections in totalcorresponding to the experiment (k=3k=3 ) in Fig. 2.. ResNet discovers that skip connections begin to demonstrate power after a few epochs compared with models without them. Interestingly, a similar phenomenon is also observed in our experiments. We term this tipping point a boundary epoch. The boundary epochs may vary from edge to edge, but are generally at the early stage. From Fig. 3, we observe that skip connections colored in red-orange progressively obtain higher architectural weights after some certain boundary epochs. Meantime, other operations are suppressed and steadily decline. We consider this benefit from the residual module as an unfair advantage by Definition 1.

Unfair Advantage. Suppose that choosing one operation among others is a competition. This competition is deemed exclusive when only restricted operations can be selected. An operation in an exclusive competition is said to have an unfair advantage if this advantage contributes more to competition than to the performance of a resulted network.

From the above discussion, we can draw Insight 1: The root cause of excessive skip connections is the inherent unfair competition. The skip connection has an unfair advantage by forming a residual module which is convenient for the supernet training, but not equally beneficial for the performance of the outcome network where the residual module is broken.

3 Non-negligible Discrepancy of Discretization

Apart from the above issue, DARTS reports that it suffers from discrepancies when discretizing continuous encodings . To verify the problem, we run DARTS in S1S_{1} on CIFAR-10, and in S2S_{2} on ImageNet. The values of softmax(α)softmax(\alpha) of the last iteration are displayed in Fig. 4 (S1S_{1}) and on the bottom left of Fig. 3 (S2S_{2}). For S1S_{1}, the largest value is about 0.3 while the smallest one is above 0.1We run DARTS 4 times and it holds every time.. This range is somewhat too narrow to differentiate ‘good’ operations from ‘bad’. For instance on edge 2 of the reduction cell, the values are very close to each other, [0.174, 0.170, 0.176, 0.112, 0.116, 0.132, 0.118], it’s hard to say that an operation weighted by 0.176 is better than the other by 0.174. For S2S_{2}, the top-1 values are not so evidently particular from layer 2 to 7. Take the second layer for example, we have to use [0.235, 0.057, 0.17, 0.016, 0.187, 0.269, 0.066] to approximate . This again confirms the existence of discrepancy.

In summary, DARTS is usually far from a good resemblance to a one-hot representation as required by its premise in Equation 3. We often have to make ambiguous choices without high confidence. Hence, we learn Insight 2: Relaxing from discrete categorical choices to continuous ones should make a close approximation.

Fair DARTS

Based on Insight 1, we propose a cooperative mechanism to eliminate the existing unfair advantage. Not only should we exploit skip connection for smoother information flow, but we also have to provide equal opportunities for other operations. In a word, they need to avoid being trapped by unfair advantage from skip connections. On this regard, we apply a sigmoid activation (σ\sigma) for each αi,j\alpha_{i,j}, so that each operation can be switched on or off independently without being suppressed. Formally, we replace Equation 2 with the following,

It’s trivial to show that even if σ(αskip)\sigma(\alpha_{skip}) saturates to 1, other operations still can be optimized cooperatively. Promising operations continue to grow their architectural weights to reduce Lval\mathcal{L}_{val}, which leads to a multi-hot approximation. Instead, DARTS attempts to derive a one-hot estimation. The difference is that we have extended the solution set. Consequently, it allows us to tackle the discretization discrepancy. We are left to find out how to drive σ(α)\sigma(\alpha) towards each extremity (0 or 1). Next, we discuss it in greater detail.

2 Resolve Discrepancy from Continuous Representation to Discrete Encoding

To abide by Insight 2, we explicitly coerce an extra loss called zero-one loss to push the sigmoid value of architectural weights towards 0 or 1. Let L0−1=f(z)L_{0-1}=f(z) denote this loss component, where z=σ(α)z=\sigma(\alpha). To achieve our goal, the loss design must meet three basic criteria, a) It needs to have a global maximum at z=0.5z=0.5 (a fair starting point) and a global minimum at 0 and 1. b) The gradient magnitude dfdz∣z≈0.5\frac{df}{dz}|_{z\approx 0.5} has to be adequately small to allow architectural weights to fluctuate, but large enough to attract zz towards 0 or 1 when they are a bit far from 0.5. c) It should be differentiable for backpropagation.

According to the first requirement, we move σ(α)\sigma(\alpha) away from 0.5 towards 0 or 1 to minimize the discretization gap. The second one enacts explicit necessary constraints. Particularly, small gradients around the peak avoid stepping easily into two ends. Larger gradients around 0 and 1 instead help to quickly capture zz nearby. Quite straightforward, we come up with a loss function to meet the above requirements, formally as,

In order to control its strength, we weight this loss by a coefficient w0−1w_{0-1}, thus the total loss for α\alpha is formulated as,

Like DARTS , the architectural weights can be optimized through backpropagation. From Equation 8, the search objective is to find an architecture of high accuracy with a good approximation from a continuous encoding to a discrete one.

Moreover, the second requirement is indispensable, otherwise the gradient-based approach may step into local minimum too early. Here we design another loss as a negative example. Let L0−1′L_{0-1}^{\prime} be the following,

It’s trivial to see that d∣z−0.5∣dz∣z>0.5=1\frac{d|z-0.5|}{dz}|_{z>0.5}=1 and d∣z−0.5∣dz∣z<0.5=−1\frac{d|z-0.5|}{dz}|_{z<0.5}=-1. Once zz stays away from 0.5, it may receive the same gradient (1 or -1) in the later iterations, thus rapidly pushing the architectural weights towards two ends. This phenomenon is illustrated in Fig. 5.

To conclude, by combining Equation 4, 6 and 8, our method which we call Fair DARTS, can be now formally written as

It is also important to recognize that our zero-one loss is specially designed for Fair DARTS. Pushing σ(α)\sigma(\alpha) of one edge towards 0 or 1 is independent of others. It cannot be directly applied to DARTS given the exclusive competition by softmax. As the architectural weights converge to their extremities, it’s natural to use a threshold value σthreshold\sigma_{threshold} in our approach to infer submodels instead of argmax .

Experiments and Results

At the search stage, we use similar hyperparameters and tricks as . We apply the first-order optimization and it takes 10 GPU hours. All experiments are done on a Tesla V100. We select our target models with σthreshold=0.85\sigma_{threshold}=0.85The maximum number of edges for a node is also limited to 2 as in DARTS.. We use the same data processing and training trick as .

Our collaborative approach performs well with skip connections aggregation. To verify this, we repetitively search 7 times on different random seeds and report the number of skip connections in Fig. 15 (see supplementary). Since the number of skip connections is more reasonable, we obtain an average top-1 accuracy 97.46%\%. Especially, the smallest FairDARTS-a reaches 97.46%97.46\% accuracy on CIFAR-10 with reduced parameters and multiply-adds. A complete result of FairDARTS searched cells are shown in the supplementary (Fig. 13, Fig. 14 and Table 8).

2 Transferring to ImageNet

As a common practice, we transfer two searched cells (FairDARTS-a and bTheir architectures are given in Fig. 13 and 14 (supplementary).) to ImageNet. We keep the same configurations and use the identical training tricks as DARTS . Compared with SNAS and DARTS, FairDARTS-A only uses 3.6M number of parameters and 417M multiply-adds to obtain 73.7%73.7\% top-1 accuracy on ImageNet validation set. FairDARTS-B also achieves state-of-the-art 75.1%75.1\% in S1S_{1} with a smaller number of parameters than comparable counterparts.

3 Searching Proxylessly on ImageNet

Relaxing exclusive competition to collaboration greatly extends the size of the search space. In ProxylessNAS , there are 19 searchable layers and each layer contains 7 choices, consisting of 7197^{19} possible models. In our approach, every choice can be activated independently, thus, S2S_{2} contains (27)19=12819({2^{7}})^{19}=128^{19} possible models. To our knowledge, this is a most gigantic search space ever proposed, about 181918^{19} times that of .

For this search phase, we train for 30 epochs with a batch size of 1024, which takes about 3 GPU days. The final architectural weight matrix (after sigmoid activation) on the bottom right of Fig. 3 is used to derive target models. Under this cooperative setting, the skip connections and other inverted bottleneck blocks can be both chosen to work together, where the former facilitates the training and the latter learn the residual information. In contrast, under the competitive setting of DARTS, it’s impossible to achieve this, as shown in the bottom left of Fig. 3. Within 19 layers have 11 skip connection operation is preferred, which cuts down the overall depth of searchable layers to 8.

To be fair, we select at most two choices per layer if there are more than two above σthreshold\sigma_{threshold} (0.75) and use the same training tricks as . We exclude squeeze and excitation and refrain from using AutoAugment tricks though they can boost the classification accuracy further. The searched model FairDARTS-D is shown in Fig. 6, which places the summation of two inverted bottleneck blocks nearby the down-sampling stage to keep more information. It also utilizes large kernels and big expansion blocks at the tail end. Further, We raise the σthreshold\sigma_{threshold} as 0.8 to get a more lightweight model FairDARTS-C. FairDARTS-C achieves 75.1%75.1\% top-1 accuracy using only 4.2 M number of parameters. To make comparisons with EfficientNetB0 , MobileNetV3 and MixNet , FairDARTS-C obtains 77.2%77.2\% top-1 accuracy with the same tricks such as squeeze-and-excitation , AutoAugment and Swish .

Ablation Study and Analysis

As unfair advantages are mainly from skip connections, if we remove them from S1S_{1} and get the reduced search space S1∖{skip}S_{1}\setminus\{skip\}, we should expect a fair play even in an exclusive competition. Several runs of this experiment also show that there is indeed no more prevailing operations that suppress others, including other parameter-less ones like max-pooling and average pooling (Fig. 7). For S1∖{skip}S_{1}\setminus\{skip\}, we run all the experiments with 7 different random seeds and we train the searched models from scratch. The best models (96.88±0.18%96.88\pm 0.18\%) are slightly higher than DARTS (96.76±0.32%96.76\pm 0.32\%)This differs from DARTS’ reported values as it trains one model for several times., but lower than FairDARTS (97.41±0.14%97.41\pm 0.14\%) in S1S_{1}. The lowered accuracy indicates that adequate skip connections are indeed beneficial for accuracy.

2 How Does Zero-One Loss Matter?

Removing Zero-One Loss. We design two comparison groups for Fair DARTS with and without zero-one loss. Other settings are kept the same. We count the distribution of the sigmoid outputs from architectural weights and plot it on the left of Fig. 8. The one without zero-one loss covers a wide range between 0 and 0.6. So we have to make ambiguous choices again. Whereas the proposed loss has narrowed the distribution into two ends around 0 and 1. To further evaluate the influences of removing L0−1L_{0-1}, we repeat the searching for 7 times using different random seeds. The averaged top-1 accuracy for these models is 97.33±\pm0.15 (532M FLOPS on average, 74 M more than FairDARTS with L0−1L_{0-1}). Therefore, although the unfair advantage is balanced, making ambiguous choices still bring noises to the final search result, which is better solved by L0−1L_{0-1}. Discrepancy elimination seems to helps find more light-weight and accurate models.

Zero-One Loss Design. We run two experiments on CIFAR-10, one with L0−1L_{0-1} (proposed) and the other L0−1′L^{\prime}_{0-1} (control). To some extent, L0−1L_{0-1} allows stepping out of the local minimum while L0−1′L_{0-1}^{\prime} selects operations only at an early stage which depends greatly on the initialization. This matches our analysis in Section 4.2. The detailed results under both loss functions are shown in Fig. 10 (supplementary).

Loss Sensitivity. As the weight w0−1w_{0-1} of this auxiliary loss goes higher, it should squeeze more operations towards two ends, but it must not overshadow the main entropy loss. We perform several experiments where an integer w0−1w_{0-1} varies within . The right of Fig. 8 shows the final number of dominant operations for each. We select a reasonable w0−1=10w_{0-1}=10 for the best trade-off.

3 Discussions From Fairness Perspective

We review the existing methods that seek to avoid the discussed weaknesses. In general, adding dropouts to operations is similar to blending them with a simple additive Gaussian noise, both reduce the performance gain from unfair advantages. Early-stopping avoids the case before unfairness prevails.

Adding dropout to skip connections reduces unfairness. The operation-level dropout inserted after skip connections by P-DARTS can be viewed as an alleviation of unfair advantage. However, it comes with two obvious drawbacks. First, this dropout rate is hard to tune. Second, it is not so effective that they must involve another prior: setting the number of skip connections in the final cell to MM. This is a very strong prior for searching good architectures .

Adding dropout to all operations also helps. Dropout troubles the training of skip connections and thus weakens the unfair advantage. Reasonably, higher dropout rates are more effective, especially for parameter-free operations. Therefore, RobustDARTS adds dropout to all operations and obtains promising results.

Early stopping matters. DARTS+ explicitly limits the maximum number of skip connections, which can be viewed as an early-stopping strategy nearby the previously mentioned boundary epoch, right before too many skip connections rise into power. RobustDARTS also exploits early-stopping when the maximal Hessian eigenvalues change too fast.

Limiting the number of skip connections is a strong prior. In the regularized search space of P-DARTS and DARTS+ , we find that simply by restricting M=2M=2, it is possible to generate competitive models even without searching. We randomly sample models from their search space and report the results in Table 3. In Experiment 1, the second group restricts the multiply-adds to be above 500M, to further leverage the average performance. Surprisingly, both groups outperform DARTS .

Random noise can break unfair advantage. Based on our theory, we can boldly postulate that adding a random noise also disrupts the unfair advantage. Therefore, on top of DARTS , we mix the skip connections’ architectural weights with a standard Gaussian noise N(0,1)\mathcal{N}(0,1), which has a cosine decay on 50 epochs. The results strongly confirm our hypothesis, as shown in Table 3. We repeat it 4 times to have similar results.

Remove unfair advantages or destroy the exclusive competition? In principle, we can break either one of the indispensable factors to avoid collapse. However, FairDARTS breaks the latter which is simple and effective. Besides, it paves the way to eliminate the discrepancy by scheming an auxiliary loss L0−1L_{0-1}. Otherwise, the discrepancy issue remains hard to solve. However, to tackle the discrepancy issue, it’s promising that the existing approaches might benefit from tricks like L0−1L_{0-1}. This remains to be our future work.

Conclusion

We unveil two indispensable factors of the DARTS’s aggregation of excessive skip connections: unfair advantages and exclusive competition. We prove that breaking any one of them can improve the robustness. First, by allowing collaborative competition, each operation develops its architectural weight independently. Meanwhile, the non-negligible discrepancy of discretization is reduced at maximum by coercing a novel auxiliary loss which polarizes the architectural weights. In this regard, we achieve state-of-the-art performance both on CIFAR-10 and ImageNet. Second, disturbing the differentiable process with a Gaussian noise removes unfair advantage which leads to competitive results.

One of our future work is to make it more memory-friendly. As Gumbel softmax is used to replace categorical distribution , is there a similar way to our approach? More methods remain to be explored on our basis.

Supplementary of “Fair DARTS: Eliminating Unfair Advantages in Differentiable Architecture Search”

Xiangxiang Chu Tianbao Zhou Bo Zhang Jixiang Li

Weight-sharing Neural Architecture Search

Weight-sharing in neural architecture search is now prominent because of its efficiency . They can roughly be divided into two categories.

One-stage approaches. Specifically, ENAS adopts a reinforced approach to train a controller to sample subnetworks from the supernet. ‘Good’ subnetworks yield high rewards so that the final policy of the controller is able to find competitive ones. Notice the controller and subnetworks are trained alternatively. DARTS and many of its variants are also a nested optimization based on weight-sharing but in a differentiable way. Besides, DARTS creates an exclusive competition by selecting only one operation on an edge, opposed to ENAS where more operations can be activated at the same time.

Two-stage approaches. There are some other weight-sharing methods who use the trained supernet as an evaluator . We need to make a distinction here as DARTS is a one-stage approach. The supernet of DARTS is meant to learn towards a single solution, where other paths are less valued (weighted). Instead like in , all paths are uniformly sampled, so to give an equal importance on selected paths. As the supernet is used for different purposes, two-stage approaches should be singled out for discussion in this paper.

Search Spaces

There are two search spaces extensively adopted in recent NAS approaches. The one DARTS proposed, we denote as S1S_{1}, is cell-based with flexible inner connections , and the other S2S_{2} is at the block level of the entire network . We use the term proxy and proxyless to differentiate whether it directly represents the backbone architecture. Our experiments cover both categories, if otherwise explicitly written, the first is searched on CIFAR-10 and the second on ImageNet.

Search Space S1S_{1}. Our first search space S1S_{1} (show in Fig. 9) follows DARTS with an exception of excluding the zero operation, as done in . Namely, S1S_{1} works in the previously mentioned DAG of N=7N=7 nodes, first two nodes in cell ck−1c_{k-1} are the outputs of last two cells ck−1c_{k-1} and ck−2c_{k-2}, four intermediate nodes with each has incoming edges from the former nodes. The output node of DAG is a concatenation of all intermediate nodes. Each edge contains 7 candidate operations:

max_pool_3x3, avg_pool_3x3, skip_connect,

Search Space S2S_{2}. The second search space S2S_{2} is similar to that of ProxylessNAS which uses MobileNetV2 as its backbone. We make several essential modifications. Specifically, there are L=19L=19 layers and each contains N=7N=7 following choices,

Inverted bottlenecks with an expansion rate xx in (3,6), a kernel size yy in (3,5,7), later referred to as ExxKyy,

Experiment Details

The list of all experiments we perform for this paper are summarized in Table 4.

To summarize, we run DARTS in S1S_{1} and S2S_{2}, confirming the aggregation of skip connections in both search spaces. We also show the large discretization gap in S2S_{2}.

In comparison, we run Fair DARTS in S1S_{1} and S2S_{2} to show their differences. First, due to our collaborative approach, we can allow a substantial number of skip connections in S1S_{1}, see Fig. 15 (supplementary). Second, Fig. 12 (supplementary) exhibit the final heatmap of architectural weights σ\sigma where skip connections coexist with other operations, meantime, the values of σ\sigma are also close to 0 or 1, which can minimize the discretization gap.

We initialize all architectural weights to 0Its sigmoid output is 0.5 (fair starting point). and set w0−1=10w_{0-1}=10 . We comply with the same hyperparameters with some exceptions: a batch size 128, learning rate 0.005, and Adam optimizer with a decay of 3e-3 and beta (0.9, 0.999). Moreover, we comply with the same strategy of grouping training and validation data as . The edge correspondence in search space S1S_{1} is given in Table 5. We also use the first-order optimization to save time.

We illustrate some of the searched cells in Fig. 13 and Fig. 14. The normal cell of FairDARTS-a is rather simple and only 3 nodes are activated, which is previously rarely seen. In the reduction cell, more edges are activated to compensate for the information loss due to down-sampling. As reduction cells are of small proportion, FairDARTS-a thus retains to be lightweight. Other searching results are listed in Table 8.

We use the SGD optimizer with an initial learning rate 0.02 (cosine decayed by epochs) for the network weights and Adam optimizer with 5e-4 for architectural weights. Besides, we set σthreshold\sigma_{threshold} = 0.75 and w0−1w_{0-1} = 1.0.

There is a minor issue need to concern. The innate skip connections are removed from inverted bottlenecks since skip connections have already been made a parallel choice in our search space.

3 Zero-one Loss Comparison

Fig. 10 compares results on two different loss designs. With the proposed loss function L0−1L_{0-1} so that Fair DARTS is less subjective to initialization. The sigmoid values reach to their optima more slowly than that of L0−1′L^{\prime}_{0-1}. It also gives a fair opportunity near 0.5, the 3×\times3 dilation convolution on Edge (3,4) first increases and then decreases, which matches the second criteria of loss design.

4 Single-level vs. Bi-level Optimization

The 5×55\times 5 separable convolution on edge (2, 2) under bi-level setting weighs higher at the early stage but much decreased in the end, which can be viewed as robustness to a local optimum. See Fig. 11.

5 Results on COCO Object Detection

We use our models as drop-in replacements of the backbone of RetinaNet . Here we only consider comparable mobile backbones. Particularly, we use the MMDetection toolbox since it provides good implementations for various objective methods. All the models are trained and evaluated on MS COCO dataset (train2017 and val2017 respectively) for 12 epochs with a batch size of 16. The initial learning rate is 0.01 and divided by 10 at epochs 8 and 11. Table 7 shows that our model achieves the best average precision 31.9%31.9\%.

Figures

Fig. 17 and 21 gives the complete softmax evolution when running DARTS on CIFAR-10 in S1S_{1}. Fig. 23 is a similar case except the skip connection is removed.

Fig. 18 gives the complete sigmoid evolution when running Fair DARTS on CIFAR-10 in S1S_{1} with L0−1′L^{\prime}_{0-1}.

Fig. 19, 20 gives the complete sigmoid evolution when running Fair DARTS on CIFAR-10 in S1S_{1} with single-level and bi-level optimization respectively. Fig. 22 is a stacked-bar version of Fig. 20.

Although our auxiliary loss L0−1L_{0-1} is particularly designed for Fair DARTS, it is interesting to see how DARTS behave under this new condition. Fig. 16 and 24 give us such an illustration, where L0−1L_{0-1} has the same weight w=10w=10 as in Fair DARTS. Not surprisingly, under the exclusive competition by softmax, skip connections exploit even more from unfair advantages as we drive the weak-performing operations towards zero (Fig. 24). Noticeably, there is a domino effect, as the weakest operation decrease its weight, the second weakest follows, so on and so forth. This effect speeds up the aggregation and as a result, more skip connections stand out (Fig. 16). As the rest better-performing operations are contending with each other, it still finds difficulty to determine which one is the best. Besides, training its inferred best models reaches 96.77 ±\pm 0.29% on CIFAR-10 (run 7 times each with different seeds), which is not too different from the original DARTS. Therefore, we conclude that applying L0−1L_{0-1} alone is not enough to solve the existing problems in DARTS. In fact, L0−1L_{0-1} cannot handle the softmax case inherently.

References