|1
2|
|3
METRICS OF
FORMS OF IMPRESSIONS explorations in image sequence transformations
Anna Maragkoudaki MAS / Architecture and Information CAAD / ETH Zurich 2014 / 15
4|
Many thanks to Ludger Hovestadt and Vera BĂźhlmann for broadening our horizons.
|5
There are no facts, only interpretations. Friedrich Nietzsche
6|
|7
| CONTENTS 11-30
experiment no.1
31-40
experiment no.11
41-64
experiment no.111
65-93
experiment no.1111
95-99
code
101-107
text
109
readings
111
credits |
8|
|9
IMPRESSIONS This project speculates on how digital and technological literacy is affecting and will continue to challenge the way we perceive visual input and our environment, renegotiating the way we exist in the digital realm. Is what we understand from these inputs something concretized or can it be seen in a different perspective? Can human visual perception of pictorial mediums, that reproduce reality, become so abstract and generic that these mediums will be represented by a geometry formalized from the matrix of their intensities? Employing aspects and techniques from medical imaging, visual perception, computer vision and lossy video compression it focuses on articulating on visual impressions and their data transformations.
10 |
| 11
experiment no.1 image sequence to volume to form
12 |
video index: 08
video index: 07
video index: 06
video index: 05
video index: 04
video index: 03
video index: 02
video index: 01
| 13
For a first set of tests, found videos from the internet are collected and imported into Mathematica. There, the sequences get manipulated in order to reveal the structure that can be found in the movement through the passing of time, exploring how this might render as a geometry and whether the output of this process is the anticipated. It is an animation whose states crystalize over time in an additive manner. No form of image manipulation is done on the images beforehand. The number of frames selected is what defines the resolution of the produced 3D Image and graphics. A fixed-interval range of frames are extracted for each clip. These get transformed into volume by arraying all the selected range of frames [Image3D]. Image manipulation and mathematical morphology transforms are applied so the outputs are clearer.
14 |
Cloud time-lapse Different colour function renderings Video index: 06 Frames 1-200 Frame count 1
| 15
Ink dissolving in water Video index: 05 In this video consequent frames are extracted, from frame no. 1 to 195 with a step of 1 frame per layer of 3D volume. This object is a depiction of the array of frames stacked in the z axis, manipulated with a colour function that allows the extraction of the background so the structure of the movement is more apparent.
16 |
Ink-drop dissolving into water Video index: 04 Frames 1-900 Frame count 1
| 17
18 |
Milk dissolving in water Video index: 01 Frames 1-195 Frame count 2
| 19
Swarm of birds Video index: 02 Frames 300-700 Frame count 1
20 |
| 21
Ink dissolving in water Video index: 07 Frames 1-700 Frame count 10
22 |
1
9
14
| 23
5
17
From left to right and top to bottom, 3D Images and their corresponding Bottom-hat transforms: 1,2-Milk dissolving in water 3,4-Swarm of birds 5,6-Fish swimming 7,8-Inkdrop in water 9,10- Ink dissolving in water 11,12,13-Cloud time-lapse 14-15,16-Waterfall 17-18-Symmetric ink dissolving in water
24 |
movie index: 08
movie index: 07
movie index: 06
movie index: 05
movie index: 04
movie index: 03
movie index: 02
movie index: 01
| 25
A second input to the above workflow, is image sequences from movies. The camera is moving through the space, informing the viewer about the spatial configuration of the plateau instead of the character or the action. The spaces, when combined, form a continuous path which can be reconfigured, in a multitude of ways, in a speculative filmic path that rearticulates the extracted impressions of the scenes. The number of possible outcomes has the capacity to be infinite, both as a figural geometry and as spatial deployment.
26 |
Movie indexes from left to right: 01-08 Different scale manipulations and colour functions.
| 27
28 |
Movie index: 01 Conversion of graphic object to geometry with high frame count
| 29
Movie index: 03 - 04
30 |
Movie index: 05-06
| 31
experiment no.2 image morphological graphs to topology to form
32 |
The mean of all the frames of each movie image sequence. From top to bottom and left to right movie indexes: 1-8
| 33
Morphological graphs of each image. Representation of the connectivity of the morphological branches of the image after application of thinning.
34 |
| 35
The edge rules of the above graphs get plotted in space. The topological information of the 3D graphs returns a 5D list of connectivity information (defined from how many neighbors are defined for each point).
36 |
The 5D topology lists, of each of the 8 indexes, are merged into a final 8D list. This resulting list is reduced dimension through a self-organizing map in which each dimension represents data from one scene. The produced point cloud instances contain all the movie indexes at the same time, but the weights differ in respect to which scene is selected to prevail over the others. Every point cloud represents a complete filmic path. Rows from left to right: movie indexes: 1-8. Columns showcase different clusterings with respect to index weights and number of iterations.
| 37
38 |
The selected instances from the above table (first row), where each represents one movie index, get reclustered through a second self-organizing map. From left to right: movie indexes: 1-8. From top to bottom: Point cloud from first SOM (p. 36-37). Re-organised points through a 80-80 neurons 3D SOM. Resulting fitted wireframe and mesh geometry.
| 39
40 |
| 41
experiment no.3 image sequence to optical flow data to form
42 |
| 43
The optical flow algorithm is run on video index: 05 from experiment I. Outputs is a set of 8-dimentional data (r, g, b, x0, y0, u, v, time) where: x0, y0 are the new macroblock positions for each frame, u, v their displacements per frame (u for x axis, v for y axis), RGB values represent the directionality of the vectors, Time is the timestamp of the movies frames.
44 |
| 45
Each point represents the macroblock’s position on each frame, starting from the x0 and y0 position adding the u, v position difference of each movement-vector (x0+u, y0+v). The z axis is represented by the sequential order of the video frames multiplied by a scalar. Point cloud of macroblock movement in 3D space Video index: 05
46 |
| 47
The quantization process is achieved through selecting the numerically largest motion vectors. This linear process allows for deviation of the final geometry analogous to the amount of detail / meaning wanted. Video index: 05
movie index: 08
movie index: 07
movie index: 06
movie index: 05
movie index: 04
movie index: 03
movie index: 02
movie index: 01
48 |
| 49
Application on the movie indexes.
50 |
Extracted voxel geometry: Raw point cloud data, x0,y0 positions of macroblocks. From left to right and top to bottom: movie indexes: 1-8
| 51
# new frame : 0 0.012495071,0.528363,0.729571,15,175,4.1771536,0.2430266,203.0,196.0,185.0,0.1458367 0.003715247,0.47660935,0.7598377,15,185,-15.683749,0.7391992,203.0,196.0,185.0,0.14607881 0.003435105,0.5101966,0.74318415,15,195,14.910049,0.30616724,203.0,196.0,185.0,0.14629248 0.75808346,0.9248344,0.15854108,15,205,8.847027,14.563201,203. 0,196.0,185.0,0.14652339 0.60753846,0.9872792,0.20259118,15,215,10.802205,48.94704,203. 0,196.0,185.0,0.14676505 0.14752069,0.85364056,0.4994194,15,225,50.55751,50.72407,203.0,196.0,185.0,0.14701228 0.01679784,0.6264397,0.6783812,15,235,91.34899,23.903332,203.0,196.0,185.0,0.14723295 0.003588289,0.4452928,0.7755594,15,245,-85.16464,9.385592,203.0,196.0,185.0,0.14743403 0.19286433,0.1066744,0.85023063,15,255,-31.873465,40.81795,203.0,196.0,185.0,0.1476486 0.42986202,0.006839037,0.7816495,15,265,-3.737833,26.281813,203.0,196.0,185.0,0.14785014 0.013846159,0.6104459,0.68785393,15,275,33.337475,7.573711,203.0,196.0,185.0,0.14805679 0.021967828,0.6412591,0.6683866,15,285,31.173002,9.211661,203.0,196.0,185.0,0.148256 0.014078915,0.59192777,0.6969967,15,295,8.899969,1.6837186,203.0,196.0,185.0,0.14844914 0.0063563287,0.52843827,0.7326027,25,175,8.914069,0.5135301,203.0,196.0,185.0,0.14868006 0.0021435618,0.47443482,0.7617108,25,185,-33.46716,1.7185559,203.0,196.0,185.0,0.14892262 0.0012752116,0.49524343,0.7517407,25,195,-39.81738,0.3797569,203.0,196.0,185.0,0.14913723 0.3647415,0.9795451,0.32785666,25,205,7.7527394,27.486542,203.0,196.0,185.0,0.14934202 0.37446913,0.9832943,0.3211183,25,215,18.762096,72.23414,203.0,196.0,185.0,0.15059412 0.097053915,0.79521585,0.55386513,25,225,83.547005,61.210175,203.0,196.0,185.0,0.1508591 0.013073683,0.61174315,0.68759155,25,235,116.93653,26.83539,203.0,196.0,185.0,0.15108909 0.003365934,0.44613528,0.77524936,25,245,-109.533394,11.879943,203.0,196.0,185.0,0.1513228 0.11535421,0.18155599,0.8515449,25,255,-59.96579,49.645016,203.0,196.0,185.0,0.15155698 0.04864919,0.28698117,0.8321848,25,265,-49.80427,23.505548,203.0,196.0,185.0,0.15178557 0.01972711,0.63718617,0.67154336,25,275,92.69182,26.476692,203.0,196.0,185.0,0.15200995 0.024295539,0.65160024,0.6620521,25,285,65.7773,20.96229,203.0,196.0,185.0,0.15223108 0.0139605105,0.6040453,0.6909971,25,295,16.482088,3.528281,203.0,196.0,185.0,0.15244709 0.3830899,0.96728766,0.32481122,35,165,0.6385196,2.5521514,203.0,196.0,185.0,0.15266074 0.18287012,0.8785943,0.4692678,35,175,-
Sample data extracted from the algorithm through a CSV file containing, for each macroblock in each frame, the position, displacement, directionality colour deviation and time information.
1
52 |
| 53
Different point cloud mappings of the data extracted from the CSV file. From left to right movie indexes: 1-8 From top to bottom: x0,+u y0+v, vector magnitude x0,+u y0+v, time x0 , x0+u, time y0 , y0+v, time x0+u, x0, time y0+v, y0, time
54 |
Data mapping on spheres for value comparison between scenes. Columns from left to right: movie indexes: 1-8 Rows from top to bottom: mappings of x, y, u, v, x0+u, y0+v, magnitude
| 55
56 |
Visualization of the point cloud mappings, preserving the grid of the macroblocks.
| 57
58 |
Visualization of the spatial deployment of one of the elements. The intensities are clearly manifested through time (z axis).
| 59
60 |
Manipulation of macroblock positions of each scene separately. First row: x0+u, y0+v, time mapping Second row: data through 6D SOM Third row: visualization of data (marching cubes).
| 61
Visualization of movie index: 04
62 |
Top: point clouds of u, v, time data, movie indexes: 1-8 Bottom: magnitude from u, v values through 8D SOM. The produced point cloud instances contain all the movie indexes simultaneously, but the weights differ in respect to which data is preferred.
| 63
64 |
| 65
experiment no.4 = = no. 1 + no. 2 + no. 3
66 |
| 67
Data source: experiment no.1+2 Combined impressions. The convex hulls of the trained point clouds of the optical flow data (p. 60, second row) are manipulated by displacing the mesh vertices according to their corresponding Image3D impressions (p. 26-27). Rows from left to right movie indexes: 1-8. Columns from bottom to top, different generations of each scene.
68 |
From left to right movie indexes 07, 02, 08
| 69
70 |
Renegotiating the materiality of the expected.
| 71
72 |
Defining the path Diagrams of camera movement of the movie scenes are formalized with respect to each rooms plan. The arrows represent the view angle as the camera moves through the filmic space.
| 73
Diagram of one possible configuration of the scenes of the film as a continuous path with respect to ‘dead-end’ spaces or intermediate spaces. This route can have different configurations.
74 |
| 75
3D representation of the topological relation of the above schema. Each space is represented by its plan. Many possible spatial configurations are possible.
76 |
| 77
Data source: experiment no.3 A mapping from each scene is chosen to reconstruct the filmic path following the above topological configuration of the filmic spatial continuity.
78 |
| 79
3D visualization of the above assembly.
80 |
| 81
Data source: experiment no.2 The selected re-clustered scene instances of the image sequence graphs is, in the same way as above, configured in the topology of the path. Left: point cloud of a configuration Right: Different configuration with line mappings.
82 |
| 83
3D visualization of the above line assembly.
84 |
Data source: experiment no.3 The magnitudes of the vector fields produced by the optical flow of the movie indexes are input in a 8D SOM where they produce different modules that contain the meaning of all scenes but each one has an inclination towards one scene, in respect to the maps weight adjustment. Configuration according to the topology schema. Top: Trained point cloud assembly and its convex hull.
| 85
86 |
| 87
3D visualization of the above assembly.
88 |
| 89
3D visualization of the above assembly expressed in wireframe geometry.
90 |
Combination of data to signify the objective character of the assumptions of visual perception and the subjective character of impressions.
| 91
92 |
| 93
Expression of the subjective and objective meaning.
94 |
| 95
code snippets
96 |
Experiment I+II
| 97
Experiment II+III
98 |
Experiment III
| 99
100 |
| 101
text
102 |
What if patterns showing affinity instead of being in succession where treated as one complex pattern and read globally? Claude Levi-Strauss, The Structural analysis of Myth (1955)
1. Turing, A. (1950). I.—Computing Machinery and Intelligence. Mind, Lix(236), pp.433-460.
Now more than ever, we are surrounded by objects, no matter how these manifest or not in the material world. The virtualization of these objects is profound and widespread. What we don’t usually acknowledge though is that these objects solidify in a single manifestation, while they can exist in many different forms and contexts. Changing the appearance of objects does not necessarily mean that they escape from their essence or content. Transformations are a way of determining what is hidden in the appearance of things. Rearticulating a problem in its own spectrum can allow us to get a different perspective of it and investigate hidden structures that can lie in its data. Since the digitization of the images of things, their potentiality in terms of manipulation has significantly expanded. Re-formulation of digital material is a common process that happens in a structured and predicted manner that preserves of the image of the ideal object intact. What happens when this transformation manifests in an unusual way? By virtually redefining an object, it expresses in ways in which it articulates a different, parallel state of its freedom. This does not minimize the importance of its initial configuration but enables it to move simultaneously through multiple strata of meaning and being. By deviating from the preprogrammed action it deviates as well from the expected impression it creates. A non-descriptive depiction of a digital visual object and its correlation with its actual image depends on human cognition and visual perception. Consequently two questions arise. How will humans be able to perceive visual stimuli in the future? What happens in the digital age, where technical and digital literacy has enabled the viewer to exercise his/hers visual perception in a degree, where abstraction is not any more a constraint? Existing in the digital age, the amount of information we receive daily is immense. Almost all of our inputs, whether these are interaction with humans or visual stimuli, come from the digital domain. This immersion into a new virtuality is composed of a vast amount of abstract objects such as situations, impressions and intentions. As suggested by the fields of cognition and perception we tend to think about these immaterial concepts as concrete objects that formalize, in our brains, with a specific materiality. Can the opposite happen as well, instead of formalizing the implicit, generalize the explicit? Maybe humans can adjust efficiently to different morphological translations of mediums, for e.g. get an impression of a movie plot while looking at a form. Now, and in the digital future, specificity will dissolve in the generic, leaving an abstract diagrammatic that will suggest forces and intensities of meaning which will not exist in realistic renderings. Human perception will be able to capture the topological differences of visual inputs, producing a rhizome that shifts through mediums and formalizations and bundles in a holistic, complete and continuous, way the abstract data rather than their direct-realistic representation. In the beginning of the 40s, the brain was thought of as a
| 103
2. Enns, J., & Lleras, A. (2008). What’s next? New evidence for prediction in human vision. Trends in Cognitive Sciences, 12, 327-333. From Zacks, J. (2016). How We Organize Our Experience into Events. [online] Apa.org. Available at: http://www. apa.org/science/about/psa/2010/04/sci-brief.aspx [Accessed 31 Aug. 2016].
3. Zacks, J.M. and Swallow, K.M. (2007) ‘Event segmentation’, Current Directions in Psychological Science, 16(2), pp. 80–84. doi: 10.1111/j.14678721.2007.00480.x. 4. Zacks, J. (2016). How We Organize Our Experience into Events. [online] Apa.org. Available at: http://www.apa.org/science/about/psa/2010/04/scibrief.aspx [Accessed 31 Aug. 2016].
computer and investigations on its function commenced1. Now, this is coming even more into play as our perception has learned to accommodate digitality and our minds adapt to the operational mode of machines, algorithms and all internetrelated input. We have already developed a mental processing scheme to operate in an indexical manner, from our social media friends list to everyday task organization. We have learned to assume that the immaterial is a form of material object, we take into account relations between such objects. This era has depleted us of specificity and has left us wandering in a cloud of implicit articulations and assumptions. Whatever captures our attention, is temporal and in a few milliseconds fades away, the next one comes, and so on so forth. How can we cope with this rhythmic alteration of scenes? Our perception does not rely explicitly on a mere translation of signals. For example, previous empirical knowledge of movement (affordances), colors and forms help predict objects2 (predictive comprehension in event segmentation theory, Enns & Lleras, 2008), movements and situations leaving a lot of space for individualized interpretations of events and recognition of possibilities of interaction. Hermann von Helmholtz claims that “the human eye is so poor optically in the degree that makes vision impossible”. He argues that vision is the result of unconscious inferences and our understanding of visual stimuli comes from making conclusions from incomplete data. So our understanding of our environment is an impression based on assumptions made by our brain. There is enough evidence on the field of cognitive science and visual perception that suggests that human event perception is much more abstract as we imagine. Contemporary studies in the field, support the Event Segmentation theory, which is a way of clustering events in image sequences. According to this theory, comprehension comes with perceiving event boundaries. These boundaries are hierarchically structured, so that fine-grained events are clustered into larger coarsegrained events3 (Zacks, Tversky, & Iyer, 2001). Our minds replace the fine-grained mosaic with an abstraction of it, where perceptual event boundaries are the ones that signify the change in meaning. This concludes in that we tend to construct relations between our perceptions, constructing a form of diagram. But this is not the only abstractness that our vision creates. Humans can encode information in a lower bit rate than the one information, from, for example, a television, is transmitted. This means that there is a segment of information that we are managing without4 (Zacks, 2010) and subsequently choose to disregard, without this being a factor of understanding less. Another example of visual abstraction is chunking, something that the Gestalt psychologists in the early 20s identified as a key factor of perception and cognition. Chunking has mostly been investigated so far mainly in terms of spatial discretization. A new body of research has shown that just as segmenting in space is important for understanding objects, segmenting in time is as important for understanding events. The above suggest, that when we are watching a movie, our brains divide the image sequences in groups. Usually this
104 |
segmentation is site-specific, it involves the location of the actors and how this space changes over time. Extracting these kind of segments from a film is like extracting words out of a narrative. These are the constituent parts of what makes the movie and can be rearranged to alter that movie. Apart from the fact that our visual perception is dealing with visual stimuli abstractly, digital media has forced us to downsample our perception. Now that we are mostly experiencing reality through the digital domain many of our visual inputs are subjected to various and different forms of compression, almost always a lossy compression. Material we see is being compressed, due to the need for online transmission, to an extent, where data we do not perceive are discarded through psychovisual redundancy layers of the compressor’s quantization, during encoding.Visual perception however allows for low quality footage filled with artefacts to be seen without any problem. But, like our perception can recognize events as abstract stimuli, when we decode visual material, in the same way this abstraction of compression can be infinite. In this context the interpretation of film is left to the spectator and sometimes this can be an unusual way of perceiving the visible. | This project explores the transformational potential of video footage by their data. It aims to deviate from a literal translation of their temporality to the z axis. In this mindset, it was not intended to conclude in forms that would depict a realistic result, directly correlated with the real image of the movie, but a more intuitive translation which would challenge the concept of the ideal, singular object. Taking movies as input, various tools are utilized, in order to convert them to structural diagrams of meaning. The explicit details of the individual scenes, morphological transforms, graphs, optical flow etc., are the tools to create a topology that will later crystalize into various figurations. The specific scenes are thought of as visual events that can be concretized in space and time, forming entities that express their intensities. A huge number of data that construct what we see is invisible to the viewer because it is hidden in the background of encoding algorithms or not visualized at all. Is there a hidden structure within these data? What is the result of their mapping in time and space? What kind of forms can we extract? Furthermore, how can these inputs be combined to actualize a movie? The experiments performed are a gradient that moves from a temporal, image-based, transformation setting to a more abstract one, involving the actual data of the material while it is being transformed. The overall question remains the same throughout: ‘How can a visual input be transformed in a diagram that solidifies the qualitative characteristics of its figuration?’ Three axioms are used as guidelines: 1: Movement creates meaning, 2: Compression is abstraction and 3: Space is a set of direction symbols. The material used is sixteen sample
| 105
5. Zacks, J. and Tversky, B. (2001). Event structure in perception and conception. Psychological Bulletin, 127(1), pp.3-21. 6. Gibson, J.J.J. (1979) The ecological approach to visual perception. Boston: Houghton Mifflin.
videos, eight with simple movements of a subject with still camera and eight with a camera that moves into the space. The movie indexes were selected in the respect that, most movies have the same abstracted sequence of events. These individual scenes, in event structure theory, are thought of as objects. Quine, in specific, describes them as objects bounded in space-time regions while Zacks and Tversky describe events as objects in the manifold of the three dimensions of space plus the one of time5. So what if a scene was represented as a concrete object with multiple dimensions? This dynamic is manifested by the temporal transformation of the image which can be represented by displaying their change in time. Is it possible though to get an understanding of the events by studying their change? According to ecological theory, perception is more about affordances6, so the perception of kinematics and dynamic variables is what weaves the potentiality of the event. The principle of kinematic specification of dynamics (KSD) states that direct perceptual qualities emerge from the dynamics of a situation.
Experiment I The first experiment looks at image sequences are an array of pictures displayed in an order defined by their timestamp. The frames of the videos get stacked in the z axis so they form a crystallization of form in time. By adding the fourth dimension, that of time, this becomes the renderer of the actualized form of the video into the three dimensional space. An inversion happens. The three dimensional reality gets captured in a two dimensional medium, into the x, y axis. Then, through the succession of the images through time the z axis is formulated, that will transform the sequence again to the three dimensional domain but this time mapping all its conditions on all times thus helping evaluate how the movement looks like. The capacity of space is increased by the passing of time. By observing the volumes from different perspectives one sees unexpected results from trivial movements that wouldn’t be correlated otherwise with the pre-image of the object depicted in the video. Experiment I I The second experiment focuses on not using time as a renderer but the image itself. The movie indexes frames are combined into one single image from which its skeleton is kept. This results in a topology, a graph, which, by extraction of its edge rules, gets mapped in 3d space. Data from the resulting 3d graphs are extracted through defining the topology that defines them. Combining the data from each scene all together through a self-organizing map creates a multitude of different instances of movies where the significance of each scene deviates in accordance to the viewers will. These permutations are mapped as a sequence of space, later on. If the skeleton of the image is its structure,
106 |
its figuration then through combining the scenes we have a figural rendering of a generic plot. Experiment III Further investigation in how 2D movement can be formalized in 3D geometry is conducted. Here the axiom that ‘compression can be seen as a form of abstraction’ is established. The material, the scenes from the movies, are compressed fairly so the detailing is lost but the overall character and intensity of the sequence remains. What we see in movies is not a perfect reality but an approximation that mainly bases itself to be understood by a collection of forces, time-related peaks, acceleration and other time related equations. Is visual abstraction also abstraction in meaning? When we move in space, our perspective of the surroundings changes, thus we perceive the environment. Accordingly in film, the movement of the camera renders the information about the plot and subsequently the viewer gets informed of the story. Structure from motion (SFM), the usual workflow for this conversion, is reproduced here in a more intuitive way, without the use of specialized SFM software but a segment of the computer vision-compression workflow, the optical flow algorithm. Optical flow is the pattern of motion objects in an image sequence when they are moving (or the camera in relation to them) and was introduced by James J. Gibson in the 1940s. It is associated with affordance, the potential for action in space, motion estimation and motion compensation procedures used in video compression standards. These compressors consist of a multitude of layers, within which, encoding and decoding is performed. Here in particular, the block-based method (Block-based DCT encoding) is applied so that each image is fragmented to pixel neighborhoods, forming a coarser pixel grid. This way, through the variation of the grid size parameter, compression / abstraction of the image is achieved. During the encoding, the footage passes from data redundancy layers where quantification is performed. In order to reduce the transmitted data, blocks of pixels, the macroblocks, are bundled together. Instead of encoding and transmitting information about the movement of the individual pixels, the displacement of the macroblocks is measured and transmitted. The optical flow algorithm is also used in computer vision due to the fact that it is possible to extract three dimensional data from two dimensional input. Since we gain knowledge about our environment though moving in it, could this displacement be the meaning, the content of the image sequence? Vector fields are extracted using the optical flow algorithm from which the output is data regarding the ‘topological mapping’ of the image sequence. This procedure allows for the intensities of the video to be capture as a qualitative characteristic of the scene that will inform the viewer, always in an abstract and intuitive way, the essence of the space or the intensity of the actions performed within it. Scenes with more or less action, thus displacement magnitude, have shorter motion vectors (the vectors resulting from the x, y movement of the macroblocks)
| 107
7. Riemenschneider, H., Donoser, M. and Bischof, H. (2016). Bag of Optical Flow Volumes for Image Sequence Recognition. Study under the doctoral program Confluence of Vision and Graphics W1209. Institute for Computer Graphics and Vision, Graz University of Technology, Graz, Austria
and signify a more subtle event taking place in space-time. At this stage this procedure is two dimensional and takes place in the plane of each frame. When the sequential passing of the frames is taken into account a third dimension is established. The data unfold in space producing a vector field mapping the change of the content over time7. Outputs from this manipulation in a set of 8-dimentional data (r, g, b, x0, y0, u, v, time). The x0, y0 values are the new macroblock positions of each frame and the u, v their displacements per frame (u for x axis, v for y axis). The RGB values do not represent the color of the actual pixels but the directionality of the vectors in the x, y axis of the frame in relation to the vector magnitude. The blue tones signify backward movement while the brighter one forward. Finally, the time is the timestamp of the movies frames. In order to reduce the dimensions of the data set from the optical flow algorithm (exported through a CSV file), they are inputted in a self-organizing map that will produce many instances of the configuration of each of these scenes based on their displacement or magnitude values. In a combinatorial way, an index of scenes is formed where permutations happen creating different articulations of the story. So we have the probability and the abstract interpretation of any film. Since what is important in the optical flow is the topological changes of the macroblocks, using a self-organizing map that protects the topology means that meaning, if we think of meaning as the topology between the macroblocks, is kept the same even though the form might change. Experiment IIII Data of all the above transformations are used by themselves or combined to produce forms which signify the change of the sequences. Some forms prevail over others and thus, make the correlation, between the raw data, more visible. In the first experiment, the convex hulls of the trained data from experiment 3 get their temporal images applied to them as textures, allowing their impressions to get manifested also into geometry and manipulate their form through displacement of their mesh vertices. In the following combinatorial experiments, the movie indexes are organized in a possible path, which can also be reconfigured in many ways. The data of each scene are mapped in space according to that initial order. The produced forms represent either structure or meaning of the film that can solidify in many diverse geometries and with different mappings, depending on which data one wants to accentuate. The results from the above procedures are certainly interesting and can produce numerous instances, or generations, of the actual images of the initial material. However, the correlation between the data before and after is not always apparent, something that is partly because of computational limitations during the recording of the data or the amount of data that can be processed simultaneously. A possible and meaningful path to investigate this further would be a study within the spectrum of mathematic transforms and frequency domains and their associated representations.
108 |
| 109
| readings Barthes, R. and Duisit, L. (1975). An Introduction to the Structural Analysis of Narrative. New Literary History, 6(2), p.237. Deleuze, G., Tomlinson, H. and Habberjam, B. (1988) One or many durations in Bergsonism. New York: Zone Books. Engeli, M. (2000). Digital stories. Basel: Birkhäuser. Film-philosophy.com. (2016). Mules on Rodowick. [online] Available at: http://www.filmphilosophy.com/vol7-2003/n56mules [Accessed 29 Aug. 2016]. Hecht, H. (2016). Film as dynamic event perception: Technological development forces realism to retreat. Universität Mainz. Hillier, B. (2016). Space is the Machine. [online] Spaceisthemachine.com. Available at: http:// spaceisthemachine.com/ [Accessed 29 Aug. 2016]. Hovestadt, L. and Bühlmann,V. (n.d.). Eigen Architecture. Keller, E. (2003). Aliquid: The analytic of the diagram, Peter Macapia, in Chronomorphology. [New York]: Columbia University in the City of New York. Knoespel, K. (2016). Diagrammatic Transformation of Architectural Space. Levi-Strauss, C. (1955). The Structural Study of Myth. The Journal of American Folklore, 68(270), p.428. Manovich, L. (1999). Database as Symbolic Form. Convergence: The International Journal of Research into New Media Technologies, 5(2), pp.80-99. Marak, L. (2016). On image compression. [online] Ujoimro design. Available at: http://www. ujoimro.com/resources/Laszlo_Marak_image_compression.pdf [Accessed 29 Aug. 2016]. Marques, O. (2016). Image Compression and Coding - Fundamentals of visual data compression, Redundancy, models, Error-free compression,Variable Length Coding (VLC). [online] Encyclopedia.jrank.org. Available at: http://encyclopedia.jrank.org/articles/pages/6760/ Image-Compression-and-Coding.html [Accessed 29 Aug. 2016]. Riemenschneider, H., Donoser, M. and Bischof, H. (2016). Bag of Optical Flow Volumes for Image Sequence Recognition. Study under the doctoral program Confluence of Vision and Graphics W1209. Institute for Computer Graphics and Vision, Graz University of Technology, Graz, Austria Rodowick, D.N. and Jameson, F. (2001) Reading the figural, or, philosophy after the new media. Edited by Stanley Fish. Durham: Duke University Press. Sultana, T. (2012). Algebra of Three Dimensional Geometric Filters and its Relevance in 3-D Image Processing. IJMA, 4(1), pp.63-73. Szeliski, R. (2011). Optical flow in Computer vision. London: Springer. Zacks, J. and Tversky, B. (2001). Event structure in perception and conception. Psychological Bulletin, 127(1), pp.3-21. |
110 |
| 111
| credits videos Video index 01: Milk dissolving in water by unknown / 2014 Video index 02: Birds swarming by M3a9m / 2009 Video index 03: Fish relaxation scene by PlayerResidentCraft / 2012 Video index 04: Ink Drip in Water by ToobStock: Free Stock Video / 2011 Video index 05: Ink & Water by joaquín / 2012 Video index 06: Moving Clouds - time lapse by NoomHDTV Video index 07: Ink drops in water by ped roschki / 2014 Video index 08: Iceland Waterfall close-up by EleB W / 2014 movies Movie index 01: Æon Flux / Karyn Kusama Movie index 02: Æon Flux / Karyn Kusama Movie index 03: 2001: A Space Odyssey / Stanley Kubrick Movie index 04: A Space Odyssey / Stanley Kubrick Movie index 05: A Space Odyssey / Stanley Kubrick Movie index 06: A Clockwork Orange / Stanley Kubrick Movie index 07: Logan’s Run / Michael Anderson Movie index 08: The Holy Mountain / Alejandro Jodorowsky code Optical Flow : Hidetoshi Shimodaira / Open Processing SOM (Grasshopper plugin) : Crow Many thanks to Constantinos Miltiadis for his help in Processing. |
112 |
| 113
114 |