0
2
0
2
r
a
M
8
]
V
C
.
s
c
[
2
v
2
4
4
1
0
.
0
1
9
1
:
v
i
X
r
a
Published as a conference paper at ICLR 2020
CLEVRER: COLLISION EVENTS FOR
VIDEO REPRESENTATION AND REASONING
Kexin Yi∗
Harvard University
Chuang Gan∗
MIT-IBM Watson AI Lab
Yunzhu Li
MIT CSAIL
Pushmeet Kohli
DeepMind
Jiajun Wu
MIT CSAIL
Antonio Torralba
MIT CSAIL
Joshua B. Tenenbaum
MIT BCS, CBMM, CSAIL
ABSTRACT
The ability to reason about temporal and causal events from videos lies at the core of
human intelligence. Most video reasoning benchmarks, however, focus on pattern
recognition from complex visual and language input, instead of on causal structure.
We study the complementary problem, exploring the temporal and causal structures
behind videos of objects with simple visual appearance. To this end, we introduce
the CoLlision Events for Video REpresentation and Reasoning (CLEVRER) dataset,
a diagnostic video dataset for systematic evaluation of computational models
on a wide range of reasoning tasks. Motivated by the theory of human causal
judgment, CLEVRER includes four types of question: descriptive (e.g., ‘what
color’), explanatory (‘what’s responsible for’), predictive (‘what will happen next’),
and counterfactual (‘what if’). We evaluate various state-of-the-art models for
visual reasoning on our benchmark. While these models thrive on the perception-
based task (descriptive), they perform poorly on the causal tasks (explanatory,
predictive and counterfactual), suggesting that a principled approach for causal
reasoning should incorporate the capability of both perceiving complex visual and
language inputs, and understanding the underlying dynamics and causal relations.
We also study an oracle model that explicitly combines these components via
symbolic representations.
1
INTRODUCTION
The ability to recognize objects and reason about their behaviors in physical events from videos
lies at the core of human cognitive development (Spelke, 2000). Humans, even young infants,
group segments into objects based on motion, and use concepts of object permanence, solidity, and
continuity to explain what has happened, infer what is about to happen, and imagine what would
happen in counterfactual situations. The problem of complex visual reasoning has been widely
studied in artificial intelligence and computer vision, driven by the introduction of various datasets on
both static images (Antol et al., 2015; Zhu et al., 2016; Hudson & Manning, 2019) and videos (Jang
et al., 2017; Tapaswi et al., 2016; Zadeh et al., 2019). However, despite the complexity and variety
of the visual context covered by these datasets, the underlying logic, temporal and causal structure
behind the reasoning process is less explored.
In this paper, we study the problem of temporal and causal reasoning in videos from a complementary
perspective: inspired by a recent visual reasoning dataset, CLEVR (Johnson et al., 2017a), we
simplify the problem of visual recognition, but emphasize the complex temporal and causal structure
behind the interacting objects. We introduce a video reasoning benchmark for this problem, drawing
inspirations from developmental psychology (Gerstenberg et al., 2015; Ullman, 2015). We also
evaluate and assess limitations of various current visual reasoning models on the benchmark.
Our benchmark, named CoLlision Events for Video REpresentation and Reasoning (CLEVRER), is a
diagnostic video dataset for temporal and causal reasoning under a fully controlled environment. The
design of CLEVRER follows two guidelines: first, the posted tasks should focus on logic reasoning in
the temporal and causal domain while staying simple and exhibiting minimal biases on visual scenes
and language; second, the dataset should be fully controlled and well-annotated in order to host the
∗indicates equal contributions. Project page: http://clevrer.csail.mit.edu/
1
Published as a conference paper at ICLR 2020
Figure 1: Sample video, questions, and answers from our CoLlision Events for Video REpresentation and Rea-
soning (CLEVRER) dataset. CLEVRER is designed to evaluate whether computational models can understand
what is in the video (I, descriptive questions), explain the cause of events (II, explanatory), predict what will
happen in the future (III, predictive), and imagine counterfactual scenarios (IV, counterfactual). In the four
images (a–d), only for visualization purposes, we apply stroboscopic imaging to reveal object motion. The
captions (e.g., ‘First collision’) are for the readers to better understand the frames, not part of the dataset.
complex reasoning tasks and provide effective diagnostics for models on those tasks. CLEVRER
includes 20,000 synthetic videos of colliding objects and more than 300,000 questions and answers
(Figure 1). We focus on four specific elements of complex logical reasoning on videos: descriptive
(e.g., ‘what color’), explanatory (‘what’s responsible for’), predictive (‘what will happen next’), and
counterfactual (‘what if’). CLEVRER comes with ground-truth motion traces and event histories
of each object in the videos. Each question is paired with a functional program representing its
underlying logic. As summarized in table 1, CLEVRER complements existing visual reasoning
benchmarks on various aspects and introduces several novel tasks.
We also present analysis of various state-of-the-art visual reasoning models on CLEVRER. While
these models perform well on descriptive questions, they lack the ability to perform causal reasoning
and struggle on the explanatory, predictive, and counterfactual questions. We therefore identify
three key elements that are essential to the task: recognition of the objects and events in the videos;
modeling the dynamics and causal relations between the objects and events; and understanding of the
Publicité
symbolic logic behind the questions. As a first-step exploration of this principle, we study an oracle
model, Neuro-Symbolic Dynamic Reasoning (NS-DR), that explicitly joins these components via a
symbolic video representation, and assess its performance and limitations.
2 RELATED WORK
Our work can be uniquely positioned in the context of three recent research directions: video
understanding, visual question answering, and physical and causal reasoning.
Video understanding. With the availability of large-scale video datasets (Caba Heilbron et al.,
2015; Kay et al., 2017), joint video and language understanding tasks have received much interest.
This includes video captioning (Guadarrama et al., 2013; Venugopalan et al., 2015; Gan et al., 2017),
localizing video segments from natural language queries (Gao et al., 2017; Hendricks et al., 2017),
and video question answering. In particular, recent papers have explored different approaches to
acquire and ground various reasoning tasks to videos. Among those, MovieQA (Tapaswi et al.,
2016), TGIF-QA (Jang et al., 2017), TVQA (Lei et al., 2018) are based on real-world videos and
human-generated questions. Social-IQ (Zadeh et al., 2019) discusses causal relations in human
social interactions based on real videos. COG (Yang et al., 2018) and MarioQA (Mun et al., 2017)
use simulated environments to generate synthetic data and controllable reasoning tasks. Compared
to them, CLEVRER focuses on the causal relations grounded in object dynamics and physical
interactions, and introduces a wide range of tasks including description, explanation, prediction and
counterfactuals. CLEVRER also emphasizes compositionality in the visual and logic context.
Visual question answering. Many benchmark tasks have been introduced in the domain of visual
question answering. The Visual Question Answering (VQA) dataset (Antol et al., 2015) marks an
important milestone towards top-down visual reasoning, based on large-scale cloud-sourced real
images and human-generated questions. The CLEVR dataset (Johnson et al., 2017a) follows a bottom-
up approach by defining the tasks under a controlled close-domain setup of synthetic images and
questions with compositional attributes and logic traces. More recently, the GQA dataset (Hudson &
Manning, 2019) applies synthetic compositional questions to real images. The VCR dataset (Zellers
2
abcdQ: What shape is the object that collides with the cyan cylinder?I. DescriptiveQ: How many metal objects are moving when the video ends?II. ExplanatoryQ: Which of the following is responsible for the gray cylinder’s colliding with the cube?a) The presence of the sphereb) The collision between the gray cylinder and the cyan cylinderA: b)III. PredictiveQ: Which event will happen nexta) The cube collides with the red objectb) The cyan cylinder collides with the red objectA: a)IV. CounterfactualQ: Without the gray object, which event will not happen?a) The cyan cylinder collides with the sphereb) The red object and the sphere collideA: a), b)A: cylinderA: 3(a) First collision(b) Cyan cube enters(c) Second collision(d) Video endsPublished as a conference paper at ICLR 2020
Dataset
Video
Diagnostic Temporal
Annotations Relation
Explanation Prediction Counterfactual
VQA (Antol et al., 2015)
CLEVR (Johnson et al., 2017a)
COG (Yang et al., 2018)
VCR (Zellers et al., 2019)
GQA (Johnson et al., 2017a)
×
×
×
×
×
(cid:88)
TGIF-QA (Jang et al., 2017)
MovieQA (Tapaswi et al., 2016) (cid:88)
(cid:88)
MarioQA (Mun et al., 2017)
(cid:88)
TVQA (Lei et al., 2018)
(cid:88)
Social-IQ (Zadeh et al., 2019)
CLEVRER (ours)
(cid:88)
×
(cid:88)
(cid:88)
(cid:88)
(cid:88)
×
×
×
×
×
(cid:88)
×
×
(cid:88)
×
×
(cid:88)
(cid:88)
(cid:88)
×
×
(cid:88)
×
×
×
(cid:88)
×
×
(cid:88)
(cid:88)
(cid:88)
(cid:88)
(cid:88)
×
×
×
×
×
×
×
×
×
×
(cid:88)
×
×
×
(cid:88)
×
×
×
×
×
×
(cid:88)
Table 1: Comparison between CLEVRER and other visual reasoning benchmarks on images and videos.
CLEVRER is a well-annotated video reasoning dataset created under a controlled environment. It introduces a
wide range of reasoning tasks including description, explanation, prediction and counterfactuals
et al., 2019) discusses explanations and hypothesis judgements based on common sense. There have
also been numerous visual reasoning models (Hudson & Manning, 2018; Santoro et al., 2017; Hu
Publicité
et al., 2017; Perez et al., 2018; Zhu et al., 2017; Mascharka et al., 2018; Suarez et al., 2018; Cao et al.,
2018; Bisk et al., 2018; Misra et al., 2018; Aditya et al., 2018). Here we briefly review a few. The
stacked attention networks (SAN) (Yang et al., 2016) introduce a hierarchical attention mechanism
for end-to-end VQA models. The MAC network (Hudson & Manning, 2018) combines visual
and language attention for compositional visual reasoning. The IEP model (Johnson et al., 2017b)
proposes to answer questions via neural program execution. The NS-VQA model (Yi et al., 2018)
disentangles perception and logic reasoning by combining an object-based abstract representation of
the image with symbolic program execution. In this work we study a complementary problem of
causal reasoning and assess the strengths and limitations of these baseline methods.
Physical and causal reasoning. Our work is also related to research on learning scene dynamics
for physical and causal reasoning (Lerer et al., 2016; Battaglia et al., 2013; Mottaghi et al., 2016;
Fragkiadaki et al., 2016; Battaglia et al., 2016; Chang et al., 2017; Agrawal et al., 2016; Finn et al.,
2016; Shao et al., 2014; Fire & Zhu, 2016; Pearl, 2009; Ye et al., 2018), either directly from im-
ages (Finn et al., 2016; Ebert et al., 2017; Watters et al., 2017; Lerer et al., 2016; Mottaghi et al., 2016;
Fragkiadaki et al., 2016), or from a symbolic, abstract representation of the environment (Battaglia
et al., 2016; Chang et al., 2017). Concurrent to our work, CoPhy (Baradel et al., 2020) studies physical
dynamics prediction in a counterfatual setting. CATER (Girdhar & Ramanan, 2020) introduces a
synthetic video dataset for temporal reasoning associated with compositional actions. CLEVRER
complements these works by incorporating dynamics modeling with compositional causal reasoning,
and grounding the reasoning tasks to the language domain.
3 THE CLEVRER DATASET
The CLEVRER dataset studies temporal and causal reasoning on videos. It is carefully designed
in a fully-controlled synthetic environment, enabling complex reasoning tasks, providing effective
diagnostics for models while simplifying video recognition and language understanding. The videos
describe motion and collisions of objects on a flat tabletop (as shown in Figure 1) simulated by a
physics engine, and are associated with the ground-truth motion traces and histories of all objects and
events. Each video comes with four types of questions generated by machine, including descriptive
(‘what color’, ‘how many’), explanatory (‘What is responsible for’), predictive (‘What will happen
next’), and counterfactual (‘what if’). Each question is paired with a functional program.
3.1 VIDEOS
CLEVRER includes 10,000 videos for training, 5,000 for validation, and 5,000 for testing. All videos
last for 5 seconds. The videos are generated by a physics engine that simulates object motion plus a
graphs engine that renders the frames. Extra examples from the dataset can be found in supplementary
material C.
Objects and events. Objects in CLEVRER videos adopt similar compositional intrinsic attributes
as in CLEVR (Johnson et al., 2017a), including three shapes (cube, sphere, and cylinder), two
materials (metal and rubber), and eight colors (gray, red, blue, green, brown, cyan, purple, and
3
Published as a conference paper at ICLR 2020
Figure 2: Sample questions and programs from CLEVRER. Left: Descriptive question. Middle and right:
multiple-choice question and choice. Each choice can pair with the question to form a joint logic trace.
yellow). All objects have the same size so no vertical bouncing occurs during collision. In each video,
we prohibit identical objects, such that each combination of the three attributes uniquely identifies one
object. Under this constraint, all intrinsic attributes for each object are sampled randomly. We further
introduce three types of events: enter, exit and collision, each of which contains a fixed number of
object participants: 2 for collision and 1 for enter and exit. The objects and events form an abstract
representation of the video. These ground-truth annotations, together with the object motion traces,
enable model diagnostics, one of the key advantages offered by a fully controlled environment.
Causal structure. Objects and events in CLEVRER videos exhibit rich causal structures. An event
can be either caused by an object if the event is the first one participated by the object, or another
event if the cause event happens right before the outcome event on the same object. For example, if a
sphere collides with a cube and then a cylinder, then the first collision and the cube jointly “cause”
the second collision. The object motion traces with complex causal structures are generated by the
following recursive process. We start with one randomly initialized moving object, and then add
another object whose initial position and velocity are set such that it will collide with the first object.
The same process is then repeated to add more objects and collisions to the scene. All collision
pairs and motion trajectories are randomly chosen. We discard simulations with repetitive collisions
between the same pair of objects.
Video generation. CLEVRER videos are generated from the simulated motion traces, including
each object’s position and pose at each time step. We use the Bullet (Coumans, 2010) physics
engine for motion simulation. Each simulation lasts for seven seconds. The motion traces are first
down-sampled to fit the frame rate of the output video (25 frames per second). Then the motion of
the first five seconds are sent to Blender (Blender Online Community, 2016) to render realistic video
frames of object motion and collision. The remaining two seconds are held-out for predictive tasks.
We further note CLEVRER adopts the same software and parameters for rendering as CLEVR.
3.2 QUESTIONS
We pair each video with machine-generated
questions for descriptive, explanatory, predic-
tive, and counterfactual reasoning. Sample ques-
tions of the four types can be found in Fig-
ure 1. Each question is paired with a functional
program executable on the video’s dynamical
scene. Unlike CLEVR (Johnson et al., 2017a),
our questions focus on the temporal and causal
aspects of the objects and events in the videos.
We exclude all questions on static object prop-
erties, which can be answered by looking at a
single frame. CLEVRER consists of 219,918 descriptive questions, 33,811 explanatory questions,
14,298 predictive questions and 37,253 counterfactual questions. Detailed distribution and split of the
questions can be found in Figure 3 and supplementary material A.
Descriptive. Descriptive questions evaluate a model’s capability to understand and reason about
a video’s dynamical content and temporal relation. The reasoning tasks are grounded to the com-
positional space of both object and event properties, including intrinsic attributes (color, material,
shape), motion, collision, and temporal order. All descriptive questions are ‘open-ended’ and can be
answered by a single word. Descriptive questions contain multiple sub-types including count, exist,
query color, query material, and query shape. Distribution of the sub-types is shown in Figure 3. We
evenly sample the answers within each sub-type to reduce answer bias.
Explanatory. Explanatory questions query the causal structure of a video by asking whether an
object or event is responsible for another event. Event A is responsible for event B if A is among
Figure 3: Distribution of CLEVRER question types.
Left: distribution of four main questions types. Right:
distribution of descriptive sub-types.
4
How many metal objects are moving when the video ends?Filter materialObjectsMetalFilter movingEventsFilter endGet frameCountWithout the gray object, which event will happen?Filter colorObjectsGrayGet counterfactsBelong toChoice programThe cube collides with the red object.CubeObjectsObjectsFilter shapeFilter colorRedFilter collisionEventsFilter collisionQuestion programCountExistQueryShapeQueryMaterialQueryColorDescriptiveExplanatoryCounter-factualPredictivePublished as a conference paper at ICLR 2020
B’s ancestors in the causal graph. Similarly, object O is responsible for event A if O participates in
A or any other event responsible for A. Explanatory questions are multiple choice questions with at
most four options, each representing an event or object in the video. Models need to select all options
that match the question’s description. There can be multiple correct options for each question. We
sample the options to balance the number of correct and wrong ones, and minimize text-only biases.
Predictive. Predictive questions test a model’s capability of predicting possible occurrences of
future events after the video ends. Similar to explanatory questions, predictive questions are multiple-
choice, whose options represent candidate events that will or will not happen. Because post-video
events are sparse, we provide two options for each predictive question to reduce bias.
Counterfactual. Counterfactual questions query the outcome of the video under certain hypothet-
ical conditions (e.g. removing one of the objects). Models need to select the events that would or
would not happen under the designated condition. There are at most four options for each question.
The numbers of correct and incorrect options are balanced. Both predictive and counterfactual
questions require knowledge of object dynamics underlying the videos and the ability to imagine and
reason about unobserved events.
Program representation.
In CLEVRER, each question is represented by a tree-structured func-
tional program, as shown in Figure 2. A program begins with a list of objects or events from the video.
The list is then passed through a sequence of filter modules, which select entries from the list and
join the tree branches to output a set of target objects and events. Finally, an output module is called
Publicité
to query a designated property of the target outputs. For multiple choice questions, each question and
option correspond to separate programs, which can be jointly executed to output a yes/no token that
indicates if the choice is correct for the question. A list of all program modules can be found in the
supplementary material B.
Question generation. Questions in CLEVRER are generated by a multi-step procedure. For each
question type, a pre-defined logic template is chosen. The logic template can be further populated by
attributes that are associated with the context of the video (i.e. the color, material, shape that identifies
the object to be queried) and then turned into natural language. We first generate an exhaustive list of
all possible questions for each video. To avoid language bias, we sample from that list to maintain
a balanced answer distribution for each question type and minimize the correlation between the
questions and answers over the entire dataset.
4 BASELINE EVALUATION
In this section, we evaluate and analyse the performances of a wide range of baseline models for video
reasoning on CLEVRER. For descriptive questions, the models treat each question as a multi-class
classification problem over all possible answers. For multiple choice questions, each question-choice
pair is treated as a binary classification problem indicating the correctness of the choice.
4.1 MODEL DETAILS
The baseline models we evaluate fall into three families: language-only models, models for video
question answering, and models for compositional visual reasoning.
Language-only models. This model family includes weak baselines that only relies on question
input to assess language biases in CLEVRER. Q-type (random) uniformly samples an answer from
the answer space or randomly select each choice for multiple-choice questions. Q-type (frequent)
chooses the most frequent answer in the training set for each question type. LSTM uses a pretrained
word embedding trained on the Google News corpus (Mikolov et al., 2013) to encode the input
question and processes the sequence with a LSTM (Hochreiter & Schmidhuber, 1997). A MLP is
then applied to the final hidden state to predict a distribution over the answers.
Video question answering. We also evaluate the following models that relies on both video and
language inputs. CNN+MLP extracts features from the input video via a convolutional neural network
(CNN) and encodes the question by taking the average of the pretrained word embeddings (Mikolov
et al., 2013). The video and language features are then jointly sent to a MLP for answer prediction.
CNN+LSTM relies on the same architecture for video feature extraction but uses the final state
of a LSTM for answer prediction. TVQA (Lei et al., 2018) introduces a multi-stream end-to-end
neural model that sets the state of the art for video question answering. We apply attribute-aware
object-centric features acquired by a video frame parser (TVQA+). We also include a recent model
that incorporates heterogeneous memory with multimodal attention work (Memory) (Fan et al.,
2019) that achieves superior performance on several datasets.
5
Published as a conference paper at ICLR 2020
Methods
Descriptive
Explanatory
Predictive
Counterfactual
per opt.
per ques.
per opt.
per ques.
per opt.
per ques.
Q-type (random)
Q-type (frequent)
LSTM
CNN+MLP
CNN+LSTM
TVQA+
Memory
IEP (V)
TbD-net (V)
MAC (V)
MAC (V+)
29.2
33.0
34.7
48.4
51.8
72.0
54.7
52.8
79.5
85.6
86.4
50.1
50.2
59.7
54.9
62.0
63.3
53.7
52.6
61.6
59.5
70.5
8.1
16.5
13.6
18.3
17.5
23.7
13.9
14.5
3.8
12.5
22.3
50.7
50.0
50.6
50.5
57.9
70.3
50.0
50.0
50.3
51.0
59.7
25.5
0.0
23.2
13.2
31.6
48.9
33.1
9.7
6.5
16.5
Publicité
42.9
50.1
50.2
53.8
55.2
61.2
53.9
54.2
53.4
56.1
54.6
63.5
10.3
1.0
3.1
9.0
14.7
4.1
7.0
3.8
4.4
13.7
25.1
Table 2: Question-answering accuracy of visual reasoning baselines on CLEVRER. All models are trained on
the full training set. The IEP (V) model and TbD-net (V) use 1000 programs to train the program generator.
Compositional visual reasoning. The CLEVR dataset (Johnson et al., 2017a) opened up a new
direction of compositional visual reasoning, which emphasizes complexity and compositionality
in the logic and visual context. We modify several best-performing models and apply them to our
video benchmark. The IEP model (Johnson et al., 2017b) applies neural program execution for
visual reasoning on images. We apply the same approach to our program-based video reasoning
task (IEP (V)) by substituting the program primitives by the ones from CLEVRER, and applying
the execution modules on the video features extracted by a convolutional LSTM (Shi et al., 2015).
TbD-net (Mascharka et al., 2018) follows a similar approach by parsing the input question into a
program, which is then assembled into a neural network that acts on the attention map over the image
features. The final attended image feature is then sent to an output layer for classification. We adopt
the same approach through spatial-temporal attention over the video feature space (TbD-net (V)).
MAC (Hudson & Manning, 2018) incorporates a joint attention mechanism on both the image feature
map and the question, which leads to strong performance on CLEVR without program supervision.
We modify the model by applying a temporal attention unit across the video frames to generate a
latent encoding for the video (MAC (V)). The video feature is then input to the MAC network to
output an answer distribution. We study an augmented approach: we construct object-aware video
features by adding the segmentation masks of all objects in the frames and labeling them by the
values of their intrinsic attributes (MAC (V+)).
Implementation details. We use a pre-trained ResNet-50 (He et al., 2016) to extract features from
the video frames. We use the 2,048-dimensional pool5 layer output for CNN-based methods, and
the 14
14 feature maps for MAC, IEP and TbD-net. The program generators of IEP and TbD-net
are trained on 1000 programs. We uniformly sample 25 frames for each video as input. Object
segmentation masks and attributes are obtained by a video parser consisted of an object detector and
an attribute network.
×
4.2 RESULTS
We summarize the performances of all baseline models in Table 2. The fact that these models achieve
different performances over the wide spectrum of tasks suggest that CLEVRER offers powerful
assessment to the models’ strength and limitations on various domains. All models are trained on the
training set until convergence, tuned on the validation set and evaluated on the test set.
Evaluation metrics. For descriptive questions, we calculate the accuracy by comparing the pre-
dicted answer token to the ground-truth. For multiple choice questions, we adopt two metrics:
per-option accuracy measures the model’s overall correctness on single options across all questions;
per-question accuracy measures the correctness of the full question, requiring all choices to be
selected correctly.
Descriptive reasoning. Descriptive questions query the content of the video from various aspects.
In order to do well on this question type, a model needs to both accurately recognize the objects
and events that happen in the video, as well as understanding the compositional logic pattern behind
the questions. In other words, descriptive questions require strong perception and logic operations
on both visual and language signals. As shown in Table 2, the LSTM baseline that relies only on
question input performs poorly on the descriptive questions, only outperforming the random baselines
6
Published as a conference paper at ICLR 2020
Figure 4: Our model includes four components: a video frame parser that generates an object-based representation
of the video frames; a question parser that turns a question into a functional program; a dynamics predictor that
extracts and predicts the dynamic scene of the video; and a symbolic program executor that runs the program on
the dynamic scene to obtain an answer.
by a small margin. This suggests that CLEVRER has very small bias on the questions. Video QA
models, including the state of the art model TVQA+ (Lei et al., 2018) achieve better performances.
But because of their limited capability of handling the compositionality in the question logic and
visual context, these models are still unable to thrive on the task. In contrast, models designed for
compositional reasoning, including TbD-net that operates on neural program execution and MAC
that introduces a joint attention mechanism, are able to achieve more competitive performances.
Causal reasoning. Results on the descriptive questions demonstrate the power of models that
combine visual and language perception with compositional logic operations. However, the causal
reasoning tasks (explanatory, predictive, counterfactual) require further understanding beyond per-
ception. Our evaluation results (Table 2) show poor performance of most baseline models on these
questions. The compositional reasoning models that performs well on the descriptive questions (MAC
(V) and TbD-net (V)) only achieve marginal improvements over the random and language-only
baselines on the causal tasks. However, we do notice a reasonable gain in performance on models
that inputs object-aware representations: TVQA+ achieves high accuracy on the predictive questions,
and MAC (V+) improves upon MAC (V) on all tasks.
We identify the following messages suggested by our evaluation results. First, object-centric repre-
sentations are essential for the causal reasoning tasks. This is supported by the large improvement on
MAC (V+) over MAC (V) after using features that are aware of object instances and attributes, as
well as the strong performance of TVQA+ on the predictive questions. Second, all baseline models
lack a component to explicitly model the dynamics of the objects and the causal relations between the
collision events. As a result, they struggle on the tasks involving unobserved scenes and in particular
performs poorly on the counterfactual tasks. The combination of object-centric representation and
dynamics modeling therefore suggests a promising direction for approaching the causal tasks.
5 NEURO-SYMBOLIC DYNAMIC REASONING
Baseline evaluations on CLEVRER have revealed two key elements that are essential to causal
reasoning: an object-centric video representation that is aware of the temporal and causal relations
between the objects and events; and a dynamics model able to predict the object dynamics under
unobserved or counterfactual scenarios. However, unifying these elements with video and language
understanding posts the following challenges: first, all the disjoint model components should operate
on a common set of representations of the video, question, dynamics and causal relations; second,
the representation should be aware of the compositional relations between the objects and events.
We draw inspirations from Yi et al. (2018) and study an oracle model that operates on a symbolic
representation to join video perception, language understanding with dynamics modeling. Our model
Neuro-Symbolic Dynamic Reasoning (NS-DR) combines neural nets for pattern recognition and
dynamics prediction, and symbolic logic for causal reasoning. As shown in Figure 4, NS-DR consists
of a video frame parser (Figure 4-I), a neural dynamics predictor (Figure 4-II), a question parser
(Figure 4-III), and a program executor (Figure 4-IV). We present details of the model components
below and in supplementary material B.
7
What shape is the second object to collide with the gray object?VideoLSTMEncoderLSTMLSTMLSTMObjectsFilter_color(gray)...Query_shape...<latexit sha1_base64="61XfzoBLfz7FPV/wQNrYIeiFZxQ=">AAAB7HicbVBNS8NAFHypX7V+VT16WSyCp5KIUI9FLx4rmLbQhrLZbtqlm03YfRFK6G/w4kERr/4gb/4bt20O2jqwMMy8Yd+bMJXCoOt+O6WNza3tnfJuZW//4PCoenzSNkmmGfdZIhPdDanhUijuo0DJu6nmNA4l74STu7nfeeLaiEQ94jTlQUxHSkSCUbSS3x8maAbVmlt3FyDrxCtIDQq0BtUvm2NZzBUySY3peW6KQU41Cib5rNLPDE8pm9AR71mqaMxNkC+WnZELqwxJlGj7FJKF+juR09iYaRzayZji2Kx6c/E/r5dhdBPkQqUZcsWWH0WZJJiQ+eVkKDRnKKeWUKaF3ZWwMdWUoe2nYkvwVk9eJ+2rumf5w3WteVvUUYYzOIdL8KABTbiHFvjAQMAzvMKbo5wX5935WI6WnCJzCn/gfP4A8WqOwg==</latexit><latexit sha1_base64="61XfzoBLfz7FPV/wQNrYIeiFZxQ=">AAAB7HicbVBNS8NAFHypX7V+VT16WSyCp5KIUI9FLx4rmLbQhrLZbtqlm03YfRFK6G/w4kERr/4gb/4bt20O2jqwMMy8Yd+bMJXCoOt+O6WNza3tnfJuZW//4PCoenzSNkmmGfdZIhPdDanhUijuo0DJu6nmNA4l74STu7nfeeLaiEQ94jTlQUxHSkSCUbSS3x8maAbVmlt3FyDrxCtIDQq0BtUvm2NZzBUySY3peW6KQU41Cib5rNLPDE8pm9AR71mqaMxNkC+WnZELqwxJlGj7FJKF+juR09iYaRzayZji2Kx6c/E/r5dhdBPkQqUZcsWWH0WZJJiQ+eVkKDRnKKeWUKaF3ZWwMdWUoe2nYkvwVk9eJ+2rumf5w3WteVvUUYYzOIdL8KABTbiHFvjAQMAzvMKbo5wX5935WI6WnCJzCn/gfP4A8WqOwg==</latexit><latexit sha1_base64="61XfzoBLfz7FPV/wQNrYIeiFZxQ=">AAAB7HicbVBNS8NAFHypX7V+VT16WSyCp5KIUI9FLx4rmLbQhrLZbtqlm03YfRFK6G/w4kERr/4gb/4bt20O2jqwMMy8Yd+bMJXCoOt+O6WNza3tnfJuZW//4PCoenzSNkmmGfdZIhPdDanhUijuo0DJu6nmNA4l74STu7nfeeLaiEQ94jTlQUxHSkSCUbSS3x8maAbVmlt3FyDrxCtIDQq0BtUvm2NZzBUySY3peW6KQU41Cib5rNLPDE8pm9AR71mqaMxNkC+WnZELqwxJlGj7FJKF+juR09iYaRzayZji2Kx6c/E/r5dhdBPkQqUZcsWWH0WZJJiQ+eVkKDRnKKeWUKaF3ZWwMdWUoe2nYkvwVk9eJ+2rumf5w3WteVvUUYYzOIdL8KABTbiHFvjAQMAzvMKbo5wX5935WI6WnCJzCn/gfP4A8WqOwg==</latexit><latexit sha1_base64="61XfzoBLfz7FPV/wQNrYIeiFZxQ=">AAAB7HicbVBNS8NAFHypX7V+VT16WSyCp5KIUI9FLx4rmLbQhrLZbtqlm03YfRFK6G/w4kERr/4gb/4bt20O2jqwMMy8Yd+bMJXCoOt+O6WNza3tnfJuZW//4PCoenzSNkmmGfdZIhPdDanhUijuo0DJu6nmNA4l74STu7nfeeLaiEQ94jTlQUxHSkSCUbSS3x8maAbVmlt3FyDrxCtIDQq0BtUvm2NZzBUySY3peW6KQU41Cib5rNLPDE8pm9AR71mqaMxNkC+WnZELqwxJlGj7FJKF+juR09iYaRzayZji2Kx6c/E/r5dhdBPkQqUZcsWWH0WZJJiQ+eVkKDRnKKeWUKaF3ZWwMdWUoe2nYkvwVk9eJ+2rumf5w3WteVvUUYYzOIdL8KABTbiHFvjAQMAzvMKbo5wX5935WI6WnCJzCn/gfP4A8WqOwg==</latexit>MaskR-CNNhOt2...t,Rt2...ti<latexit sha1_base64="9RhpuQU18Y7yXVTfNFvv13Qe5z8=">AAACF3icbVC7TsMwFHV4lvIKMLJYVEgMUCUVEowVLGwURB9SU1WO67RWHSeyb5CqqH/Bwq+wMIAQK2z8DW6agbYcydLxOfde+x4/FlyD4/xYS8srq2vrhY3i5tb2zq69t9/QUaIoq9NIRKrlE80El6wOHARrxYqR0Bes6Q+vJ37zkSnNI/kAo5h1QtKXPOCUgJG6dtkTRPYFw7fdFM4qXi8CjWF8iu9n7thTWVnXLjllJwNeJG5OSihHrWt/mxE0CZkEKojWbdeJoZMSBZwKNi56iWYxoUPSZ21DJQmZ7qTZXmN8bJQeDiJljgScqX87UhJqPQp9UxkSGOh5byL+57UTCC47KZdxAkzS6UNBIjBEeBIS7nHFKIiRIYQqbv6K6YAoQsFEWTQhuPMrL5JGpewafndeql7lcRTQITpCJ8hFF6iKblAN1RFFT+gFvaF369l6tT6sz2npkpX3HKAZWF+/mMie8A==</latexit><latexit sha1_base64="9RhpuQU18Y7yXVTfNFvv13Qe5z8=">AAACF3icbVC7TsMwFHV4lvIKMLJYVEgMUCUVEowVLGwURB9SU1WO67RWHSeyb5CqqH/Bwq+wMIAQK2z8DW6agbYcydLxOfde+x4/FlyD4/xYS8srq2vrhY3i5tb2zq69t9/QUaIoq9NIRKrlE80El6wOHARrxYqR0Bes6Q+vJ37zkSnNI/kAo5h1QtKXPOCUgJG6dtkTRPYFw7fdFM4qXi8CjWF8iu9n7thTWVnXLjllJwNeJG5OSihHrWt/mxE0CZkEKojWbdeJoZMSBZwKNi56iWYxoUPSZ21DJQmZ7qTZXmN8bJQeDiJljgScqX87UhJqPQp9UxkSGOh5byL+57UTCC47KZdxAkzS6UNBIjBEeBIS7nHFKIiRIYQqbv6K6YAoQsFEWTQhuPMrL5JGpewafndeql7lcRTQITpCJ8hFF6iKblAN1RFFT+gFvaF369l6tT6sz2npkpX3HKAZWF+/mMie8A==</latexit><latexit sha1_base64="9RhpuQU18Y7yXVTfNFvv13Qe5z8=">AAACF3icbVC7TsMwFHV4lvIKMLJYVEgMUCUVEowVLGwURB9SU1WO67RWHSeyb5CqqH/Bwq+wMIAQK2z8DW6agbYcydLxOfde+x4/FlyD4/xYS8srq2vrhY3i5tb2zq69t9/QUaIoq9NIRKrlE80El6wOHARrxYqR0Bes6Q+vJ37zkSnNI/kAo5h1QtKXPOCUgJG6dtkTRPYFw7fdFM4qXi8CjWF8iu9n7thTWVnXLjllJwNeJG5OSihHrWt/mxE0CZkEKojWbdeJoZMSBZwKNi56iWYxoUPSZ21DJQmZ7qTZXmN8bJQeDiJljgScqX87UhJqPQp9UxkSGOh5byL+57UTCC47KZdxAkzS6UNBIjBEeBIS7nHFKIiRIYQqbv6K6YAoQsFEWTQhuPMrL5JGpewafndeql7lcRTQITpCJ8hFF6iKblAN1RFFT+gFvaF369l6tT6sz2npkpX3HKAZWF+/mMie8A==</latexit><latexit sha1_base64="9RhpuQU18Y7yXVTfNFvv13Qe5z8=">AAACF3icbVC7TsMwFHV4lvIKMLJYVEgMUCUVEowVLGwURB9SU1WO67RWHSeyb5CqqH/Bwq+wMIAQK2z8DW6agbYcydLxOfde+x4/FlyD4/xYS8srq2vrhY3i5tb2zq69t9/QUaIoq9NIRKrlE80El6wOHARrxYqR0Bes6Q+vJ37zkSnNI/kAo5h1QtKXPOCUgJG6dtkTRPYFw7fdFM4qXi8CjWF8iu9n7thTWVnXLjllJwNeJG5...