<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.3.2">Jekyll</generator><link href="http://noelnegash.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="http://noelnegash.github.io/" rel="alternate" type="text/html" /><updated>2023-11-08T14:46:20+00:00</updated><id>http://noelnegash.github.io/feed.xml</id><title type="html">Code, Canvas &amp;amp; Crescendo</title><subtitle>Hi, my name is Noel Alemayehu and I am a software developer based in Addis Ababa, Ethiopia  specializing in backend, graphics and game development.  This blog is a place where I can put some of my random articles and experiments.</subtitle><entry><title type="html">Verdant</title><link href="http://noelnegash.github.io/games/verdant/2022/07/11/verdant.html" rel="alternate" type="text/html" title="Verdant" /><published>2022-07-11T06:01:27+00:00</published><updated>2022-07-11T06:01:27+00:00</updated><id>http://noelnegash.github.io/games/verdant/2022/07/11/verdant</id><content type="html" xml:base="http://noelnegash.github.io/games/verdant/2022/07/11/verdant.html"><![CDATA[<p>Game Jam</p>

<p>Game Jam</p>

<p>Game Jam</p>

<p>Game Jam
<script src="https://cdn.jsdelivr.net/npm/p5@1.6.0/lib/p5.min.js"></script>
<script src="https://cdn.jsdelivr.net/npm/p5@1.6.0/lib/addons/p5.sound.min.js"></script></p>
<link rel="stylesheet" type="text/css" href="/assets/verdant/style.css" />

<div id="verdantCanvas"></div>
<script src="/assets/verdant/sketch.js"></script>]]></content><author><name>Noel Alemayehu</name></author><category term="games" /><category term="verdant" /><summary type="html"><![CDATA[Global Game Jam entry]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://via.placeholder.com/1200x800" /><media:content medium="image" url="https://via.placeholder.com/1200x800" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Schopenhauer &amp;amp; Object Oriented Programming</title><link href="http://noelnegash.github.io/schopenhauer/programming/2020/09/19/schopenhauer.html" rel="alternate" type="text/html" title="Schopenhauer &amp;amp; Object Oriented Programming" /><published>2020-09-19T11:04:21+00:00</published><updated>2020-09-19T11:04:21+00:00</updated><id>http://noelnegash.github.io/schopenhauer/programming/2020/09/19/schopenhauer</id><content type="html" xml:base="http://noelnegash.github.io/schopenhauer/programming/2020/09/19/schopenhauer.html"><![CDATA[<style> .option-button, .reset-button, .change-button {
    border: none;
    color: white;
    text-align: center;
    text-decoration: none;
    display: inline-block;
    font-size: 16px;
    margin: 4px 2px;
    cursor: pointer;
    position: relative;
    font-family: "Montserrat", sans-serif;
    font-size: 14px;
    line-height: 24px;
    font-weight: 700;
    padding: 12px 35px;
    color: #fff;
    text-transform: uppercase;
    border-radius: 3px;
}
.reset-button {
    background-color: #4CAF50; 
    padding: 15px 32px;
}
.change-button {
    background-color: #4274d6;
    display: block;
    padding: 7px 15px;
    }
.option-button {
    background-color: #a12895;
    display: block;
    padding: 7px 15px;
}
.flex-container {
    display: flex;
    flex-direction: row;
    align-items: center;    
    justify-content: center;
}
.flex-container > * {
    margin: 30px;
}
pre {
   text-align: left !important;
}

h4, h3 {
    font-weight: bold;
    margin: 10px auto;
} </style>

<p>2020 has been a wild ride to say the least. January alone started with a million Hong Kong protesters on a New Year’s Day march, the Australian bush catching fire(ironic considering UN declared it the International Year of Plant Health), the assassination of an Iranian major general by a US drone, prompting speculations about a new war, and if that wasn’t enough, COVID-19 came along, as of now with 28.8 million infections, 900k+ dead, mass lockdowns, and the largest economic recession since the Great Depression. Honestly, so much bad stuff has happened this year that just writing it would probably trigger someone’s PTSD. Instead I leave you with this <a href="https://www.boredpanda.com/2020-year-recap">only slightly depressing recap</a> &amp; video. After all, images are worth a thousand words.</p>

<div class="flex-container"> <iframe width="877" height="493" src="https://www.youtube.com/embed/NLL7ZXXL7zc" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen=""></iframe></div>

<p>Maybe prompted by this confusion at the state of the world, I have been thinking a lot about philosophy these days. Out of the many people to write on the subject, there is one that feels somewhat fitting for our situation, Arthur Schopenhauer.</p>

<p>A notoriously pessimistic philosopher, Schopenhauer formed his system of thought from the teachings of Plato, Kant and the Indian Upanishads. He drew from the first two the idea that what we experience as the world is not actual reality and supplemented it with Indian concept of Maya, a veil that hides the true nature of the world from us. His idea of the world as will and representation is a metaphysical system that considers Plato’s “Ideas” and Kant’s “things-in-themselves” to be the same thing, an underlying structure to existence that we cannot see directly, but rather only their representations. But his idea of the will took it one step further. He said there was only one will, an undivided entity that drove all things in the universe through ceaseless striving towards its goal. The laws of physics, the movements of animals and the consciousness of human beings are all the same will manifesting itself in different gradations of development. It drives everything from gravitational to sexual attraction. You can see why this metaphysical system appealed to many thinkers that came after him, such as the philosophers Wittgenstein and Nietzsche, the scientists Schrodinger and Einstein, and the psycho-analysts Freud and Jung.</p>

<p>He was also strongly influenced by Buddhist thought, claiming that all human suffering came from attachment to desires. To Schopenhauer, desire was pain that persisted until it was satisfied. Satisfaction itself was simply the neutral lack of desire as opposed to a positive value such as joy. In essence, Schopenhauer thought that life spent in pursuit of desire was constant pain, with only short periods of boredom as reprieve, a negative-sum game. Another striking idea of his is that we live in the worst possible <em>sustainable</em> world. To Schopenhauer, the world we live in could be worse, but if that was the case, it would be so unstable it would fall into chaos. In the context of 2020, I think we can understand what he meant.</p>

<p>His ideas on aesthetics led him to be called the artist’s philosopher because they were a source of inspiration to many composers and authors. In Schopenhauer’s philosophy, the only tenable way to achieve happiness is through asceticism. Only by renouncing worldly desires and trying to understand the will itself can you transcend the cycle of pain and boredom. One of the ways he suggested it was through art. He considered art an effort to capture and express the true essence of its subject, bypassing representation and getting closer to the will. He even said music was will itself, drawing parallels between the strivings of desire and the motions of a melody.</p>

<p>In the end though, he was a bitter man, with strained relationships in his familial, social and professional circles. His relationship with his mother, who in his lifetime was a more successful writer and was usually in the company of admired contemporaries such as J. W. von Goethe, whom Schopenhauer considered a hero, was not the best. Even when he managed to gain the admiration of Goethe, he almost immediately ruined it by publishing a treatise on color theory, the same subject that they discussed with each other, that differed in many points to Goethe’s. That, along with his headstrong and self-assured attitude, led Goethe to distance himself from Schopenhauer. Even in academics, he wasn’t very popular. At one point, when he was appointed to an academic position in the University of Berlin, he scheduled his lectures to be at the same time as his intellectual nemesis, G. W. F. Hegel. Almost no one attended and he left the post two years later. His mother said it best in a letter she sent him</p>

<blockquote>
  <p>“You are unbearable and burdensome, and very hard to live with; all your good qualities are overshadowed by your conceit, and made useless to the world simply because you cannot restrain your propensity to pick holes in other people.”</p>
</blockquote>

<p>As compelling as his pessimism is, especially in 2020, over the past couple of days I was more interested in his metaphysical system of will and representation, specifically how similar it sounded to Object Oriented Programming. OOP is a computer programming model that puts emphasis on data in a program, as opposed to operations. For example, let’s say you wanted to make a car racing game. If you were to make that game using OOP standards, you would create a “class” called Car. A class is essentially an abstract schematic of what you want every car to be. It would describe the data necessary to create any model of car, such as its manufacturer and horsepower, as well as its functions, such as acceleration, steering and braking. After defining a class, all you would need to do is say <code class="language-plaintext highlighter-rouge">new Car()</code>, and pass in the necessary information to get a Car object. This structure is useful to avoid redundancy and points of failure in a program, as well as improving readability. But it is also eerily similar to Schopenhauer-ian thought. The class could be the will and all the objects you create from it could be manifestations, or representations, of that will. In a more visual example, consider the two boxes below. The first is the world as representation. It could be a game, the visual rendering of the will, or your perception of reality. The second is the world as will. It is the code of the game or the will of the universe you live in. The circle in the center is the subject and you can “simulate” its will by dragging and releasing it.</p>

<div class="flex-container">

<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
</pre></td><td class="code"><pre><span class="kd">class</span> <span class="nc">Circle</span> <span class="p">{</span>
    <span class="nf">constructor</span><span class="p">(</span><span class="nx">x</span><span class="p">,</span><span class="nx">y</span><span class="p">)</span> <span class="p">{</span>
        <span class="k">this</span><span class="p">.</span><span class="nx">position</span> <span class="o">=</span> <span class="p">[</span><span class="nx">x</span><span class="p">,</span><span class="nx">y</span><span class="p">]</span>
        <span class="k">this</span><span class="p">.</span><span class="nx">velocity</span> <span class="o">=</span> <span class="p">[</span><span class="mi">0</span><span class="p">,</span><span class="mi">0</span><span class="p">]</span>
    <span class="p">}</span>
    <span class="nf">update</span><span class="p">()</span> <span class="p">{</span>
        <span class="k">this</span><span class="p">.</span><span class="nx">position</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span> <span class="o">+=</span> <span class="k">this</span><span class="p">.</span><span class="nx">velocity</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span>
        <span class="k">this</span><span class="p">.</span><span class="nx">position</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span> <span class="o">+=</span> <span class="k">this</span><span class="p">.</span><span class="nx">velocity</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span>
        <span class="k">if </span><span class="p">(</span><span class="k">this</span><span class="p">.</span><span class="nf">hitsUniverseBoundary</span><span class="p">())</span>
            <span class="k">this</span><span class="p">.</span><span class="nf">bounceOff</span><span class="p">()</span>
    <span class="p">}</span>
<span class="p">}</span>

<span class="kd">var</span> <span class="nx">universe</span> <span class="o">=</span> <span class="p">[]</span>
<span class="nx">universe</span><span class="p">.</span><span class="nf">push</span><span class="p">(</span><span class="k">new</span> <span class="nc">Circle</span><span class="p">(</span><span class="mi">200</span><span class="p">,</span><span class="mi">200</span><span class="p">))</span>
</pre></td></tr></tbody></table></code></pre></figure>

<div id="canvas-container-0"></div>
</div>

<p>That circle must be filled with existential dread from being alone in such a small universe. Let’s give him a friend. Notice that they are both subject to the same “will” despite being individual objects with different starting positions. If we make the will dictate how they handle interactions with each other, we will see similar behavior from both.</p>

<div class="flex-container">
<div>

<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
</pre></td><td class="code"><pre> 
    <span class="kd">class</span> <span class="nc">Circle</span><span class="p">{</span>
        <span class="nf">constructor</span><span class="p">(</span><span class="nx">x</span><span class="p">,</span><span class="nx">y</span><span class="p">)</span> <span class="p">{</span>
            <span class="k">this</span><span class="p">.</span><span class="nx">position</span> <span class="o">=</span> <span class="p">[</span><span class="nx">x</span><span class="p">,</span><span class="nx">y</span><span class="p">]</span>
            <span class="k">this</span><span class="p">.</span><span class="nx">velocity</span> <span class="o">=</span> <span class="p">[</span><span class="mi">0</span><span class="p">,</span><span class="mi">0</span><span class="p">]</span>
        <span class="p">}</span>
        <span class="nf">update</span><span class="p">()</span> <span class="p">{</span>
            <span class="k">this</span><span class="p">.</span><span class="nx">position</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span> <span class="o">+=</span> <span class="k">this</span><span class="p">.</span><span class="nx">velocity</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span>
            <span class="k">this</span><span class="p">.</span><span class="nx">position</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span> <span class="o">+=</span> <span class="k">this</span><span class="p">.</span><span class="nx">velocity</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span>
            <span class="k">if </span><span class="p">(</span><span class="k">this</span><span class="p">.</span><span class="nf">hitsUniverseBoundary</span><span class="p">())</span>
                <span class="k">this</span><span class="p">.</span><span class="nf">bounceOff</span><span class="p">()</span>
            <span class="c1">// added this clause</span>
            <span class="k">if </span><span class="p">(</span><span class="k">this</span><span class="p">.</span><span class="nf">hitsOtherCircle</span><span class="p">())</span> <span class="p">{</span>
                <span class="k">this</span><span class="p">.</span><span class="nf">bounceOff</span><span class="p">()</span>
                <span class="k">this</span><span class="p">.</span><span class="nf">glow</span><span class="p">()</span>
            <span class="p">}</span>
        <span class="p">}</span>
    <span class="p">}</span>

    <span class="kd">var</span> <span class="nx">universe</span> <span class="o">=</span> <span class="p">[]</span>
    <span class="nx">universe</span><span class="p">.</span><span class="nf">push</span><span class="p">(</span><span class="k">new</span> <span class="nc">Circle</span><span class="p">(</span><span class="mi">100</span><span class="p">,</span><span class="mi">200</span><span class="p">))</span>
    <span class="nx">universe</span><span class="p">.</span><span class="nf">push</span><span class="p">(</span><span class="k">new</span> <span class="nc">Circle</span><span class="p">(</span><span class="mi">300</span><span class="p">,</span><span class="mi">200</span><span class="p">))</span>
</pre></td></tr></tbody></table></code></pre></figure>

</div>
<div id="canvas-container-1"></div>
</div>

<p>Another useful aspect of OOP is inheritance. Inheritance is an idea that since different objects can have many attributes and functions in common, we can create a class that contains them, and limit the classes of the objects themselves to describe only the specific differences, saving us from lots of redundancy. Again, eerily similar to the concept of the will coming in gradations. We can apply both to humans and animals. Humans, for all practical purposes, are animals. We eat, sleep, propagate and react to our environment in ways that let us do so for longer. But we also have the capacity for intelligent thought. In OOP terms, we are an Animal class with the added function of intelligence, and in Schopenhauer-ian thinking, we are a manifestation of the will one level above animal that overrides some basic strivings with more complex behavior that is nonetheless subservient to the will.</p>

<p>In the code below, we have created a hierarchy of will that starts with the class Circle from above, which Animal “inherits” or “extends” from, by adding the properties of lifespan and the function of reproduction, which in turn is extended by the Human class, with the property of “intelligence”, a somewhat weak aversion to being moved by irrationality, in this case dragging and releasing. Humans also have a desire to chase animals (using the same formula as gravitational force) and deal them damage on collision. A bit crude, but it does serve as a good illustration of how intelligence doesn’t necessarily mean defying the will or being above it. It is simply a higher level of the hierarchy. You can even control how humans move by dragging an animal around their vicinity. Don’t get them too close though; you might end up creating a gravitational singularity that sends the universe into utter chaos. Just saying.</p>

<div class="flex-container">
<div>

<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
</pre></td><td class="code"><pre> 
    <span class="kd">class</span> <span class="nc">Animal</span> <span class="kd">extends</span> <span class="nc">Circle</span> <span class="p">{</span>
        <span class="nf">constructor</span><span class="p">(</span><span class="nx">x</span><span class="p">,</span> <span class="nx">y</span><span class="p">)</span> <span class="p">{</span>
            <span class="k">super</span><span class="p">(</span><span class="nx">x</span><span class="p">,</span><span class="nx">y</span><span class="p">)</span>
            <span class="k">this</span><span class="p">.</span><span class="nx">life</span> <span class="o">=</span> <span class="mi">50</span>
        <span class="p">}</span>
        <span class="nf">update</span><span class="p">()</span> <span class="p">{</span>
            <span class="k">super</span><span class="p">.</span><span class="nf">update</span><span class="p">()</span>
            <span class="k">if</span><span class="p">(</span><span class="k">this</span><span class="p">.</span><span class="nf">isAtHalfLife</span><span class="p">())</span>
                <span class="k">this</span><span class="p">.</span><span class="nf">reproduce</span><span class="p">()</span>
        <span class="p">}</span>
    <span class="p">}</span>
    <span class="kd">class</span> <span class="nc">Human</span> <span class="kd">extends</span> <span class="nc">Animal</span> <span class="p">{</span>
        <span class="nf">constructor</span><span class="p">(</span><span class="nx">x</span><span class="p">,</span> <span class="nx">y</span><span class="p">)</span> <span class="p">{</span>
            <span class="k">super</span><span class="p">(</span><span class="nx">x</span><span class="p">,</span><span class="nx">y</span><span class="p">)</span>
            <span class="k">this</span><span class="p">.</span><span class="nx">life</span> <span class="o">=</span> <span class="mi">100</span>
        <span class="p">}</span>
        <span class="nf">update</span><span class="p">()</span> <span class="p">{</span>
            <span class="k">super</span><span class="p">.</span><span class="nf">update</span><span class="p">()</span>
            <span class="k">this</span><span class="p">.</span><span class="nf">chaseAnimals</span><span class="p">()</span>
            <span class="k">if</span><span class="p">(</span><span class="k">this</span><span class="p">.</span><span class="nf">hitsAnimal</span><span class="p">())</span>
                <span class="k">this</span><span class="p">.</span><span class="nf">hurtAnimal</span><span class="p">()</span>
        <span class="p">}</span>
    <span class="p">}</span>

    <span class="kd">var</span> <span class="nx">universe</span> <span class="o">=</span> <span class="p">[]</span>
    <span class="nx">universe</span><span class="p">.</span><span class="nf">push</span><span class="p">(</span><span class="k">new</span> <span class="nc">Animal</span><span class="p">(</span><span class="mi">100</span><span class="p">,</span><span class="mi">200</span><span class="p">))</span>
    <span class="nx">universe</span><span class="p">.</span><span class="nf">push</span><span class="p">(</span><span class="k">new</span> <span class="nc">Human</span><span class="p">(</span><span class="mi">300</span><span class="p">,</span><span class="mi">200</span><span class="p">))</span>
</pre></td></tr></tbody></table></code></pre></figure>

</div>
<div id="canvas-container-2"></div>
</div>

<p>You can see the “worst possible sustainable world” idea here as well. If animals couldn’t propagate, or were killed before they did, they would be wiped out in a minute and we’d be back to the cold, lonely and empty universe. But when we ensure there is always an animal alive in the code(it’s no Fibonacci, but it works), we allow the cycle to go on indefinitely, creating Schopenhauer’s universe.</p>

<div class="flex-container"><div id="canvas-container-3"></div></div>

<p>There are a lot more comparisons I can make, from public and private variables to abstraction and polymorphism. I could probably write a whole essay on the analogues between death and garbage collection, but I digress. A more interesting take is that Schopenhauer isn’t necessarily similar to OOP, but simulation theory in general. Sure enough, a short google search of “Schopenhauer simulation” gives a short quote to the same effect in <a href="https://www.thesmartset.com/schopenhauer-for-millennials/">this really good article</a>. I also found a whole paper called <a href="https://download.uni-mainz.de/fb05-philosophie-schopenhauer/files/2020/01/2006_Segala.pdf">The World as Will and the Matrix as Representation</a> which explores the same ideas through the characters and events of the classic 1999 movie. While it’s not a lot of sources, it’s fun to see other people have made the same connection.</p>

<p>So instead, let’s stick to philosophy and try to make sense of this world. The first and foremost problem is free will. How can an individual have it if will itself is an unchanging entity that lies on a cosmic scale? When you talk, are you saying what you think or what you are told to say? Probably the latter.</p>

<p>Even doing the opposite of what the will tells you is a weak freedom at best, because it is simply another form of subservience. Rejection of the will and existing independently of it are two completely different things. In the words of the great Albert Einstein:</p>

<blockquote>
  <p>“Man can do what he wills, but cannot will what he wills.”</p>
</blockquote>

<p>Or can we? This is just random speculation on my part, but what if will isn’t constant? There is no sufficient reason to say it is. I mean, sure the universe we live in has probably been operating under the same rules since before the Big Bang. But maybe that’s because it hasn’t reached a level of complexity where the will begins to change itself. What if the more we understand the will, be it through art, spirituality or science, the closer we get to changing it? After all, if someone feels a desire to change the will and their desires are decided by that same will, doesn’t that imply a process of self-induced change?</p>

<p>If this theory sounds too idealistic, let’s look at it in programming terms. Can an object change its own source code, its class? Can we make a meta-program that can edit itself? Impossible, right? Well, maybe for some languages, but there are entire families out there with explicit support for it, like Lisp, one of the first programming languages used for artificial intelligence research. Coincidence? I think not.</p>

<p>We don’t even need to go to such an obscure language for meta-programming support. Right here in your browser, we have JavaScript. It is the same language that’s running the buttons and the simulations. And it has two properties which make it very easy to simulate meta-programming, the <code class="language-plaintext highlighter-rouge">eval</code> function and <code class="language-plaintext highlighter-rouge">prototype</code> object. <code class="language-plaintext highlighter-rouge">eval</code> lets you run any text you want as code and <code class="language-plaintext highlighter-rouge">prototype</code> is basically the class itself laid bare for you to mess around with. Using these two, we can let our humans change the entire universe.</p>

<p>First, let’s get rid of the animals; there has been enough suffering. Note that we leave the class in the code because it is a dependency for human existence. I will leave it up to you to think about the implications. Secondly, let’s simulate total ignorance on the part of humans, at least initially. They will bumble around the world like hapless idiots until they interact with each other. These interactions can be communication, meditation or research, at which point the human’s “mage level” will increase, unlocking secret runes and spells that can twist the world the way they see fit. These spells which will also be triggered randomly when mages bump into each other. In the box below, we have levels 0 to 4, with powers that go from the gory conjuring of blood to the phantasmagorical changing of the sky. Be warned, the arcane magic they discover could trigger epilepsy.</p>

<div class="flex-container">
<div>

<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
</pre></td><td class="code"><pre> 
    <span class="kd">var</span> <span class="nx">spells</span> <span class="o">=</span> <span class="p">[</span>
        <span class="dl">"</span><span class="s2">spewBlood()</span><span class="dl">"</span><span class="p">,</span>
        <span class="dl">"</span><span class="s2">Mage.prototype.size = randomNumber(20,40)</span><span class="dl">"</span><span class="p">,</span>
        <span class="dl">"</span><span class="s2">Mage.prototype.color = randomColor()</span><span class="dl">"</span><span class="p">,</span>
        <span class="dl">"</span><span class="s2">Universe.prototype.gravity = !Universe.prototype.gravity</span><span class="dl">"</span><span class="p">,</span>
        <span class="dl">"</span><span class="s2">Universe.prototype.background = randomColor()</span><span class="dl">"</span>
    <span class="p">]</span>
    <span class="kd">class</span> <span class="nc">Mage</span> <span class="kd">extends</span> <span class="nc">Human</span> <span class="p">{</span>
        <span class="nf">constructor</span><span class="p">(</span><span class="nx">x</span><span class="p">,</span><span class="nx">y</span><span class="p">)</span> <span class="p">{</span>
            <span class="k">super</span><span class="p">(</span><span class="nx">x</span><span class="p">,</span><span class="nx">y</span><span class="p">)</span>
            <span class="k">this</span><span class="p">.</span><span class="nx">mageLevel</span> <span class="o">=</span> <span class="mi">0</span>
        <span class="p">}</span>
        <span class="nf">update</span><span class="p">()</span> <span class="p">{</span>
            <span class="k">if</span><span class="p">(</span><span class="k">this</span><span class="p">.</span><span class="nf">hitMage</span><span class="p">())</span> <span class="p">{</span>
                <span class="k">if</span><span class="p">(</span><span class="k">this</span><span class="p">.</span><span class="nx">mageLevel</span> <span class="o">==</span> <span class="mi">4</span> <span class="nx">and</span> <span class="nb">Math</span><span class="p">.</span><span class="nf">random</span><span class="p">()</span> <span class="o">&gt;</span> <span class="mf">0.3</span><span class="p">)</span>
                    <span class="k">this</span><span class="p">.</span><span class="nx">mageLevel</span><span class="o">++</span>
                <span class="k">else</span> <span class="k">if </span><span class="p">(</span><span class="nb">Math</span><span class="p">.</span><span class="nf">random</span><span class="p">()</span> <span class="o">&gt;</span> <span class="mf">0.3</span><span class="p">)</span>
                    <span class="nf">eval</span><span class="p">(</span><span class="nx">spells</span><span class="p">[</span><span class="nf">floor</span><span class="p">(</span><span class="nf">random</span><span class="p">(</span><span class="k">this</span><span class="p">.</span><span class="nx">mageLevel</span><span class="p">))])</span>
            <span class="p">}</span>
        <span class="p">}</span>
    <span class="p">}</span>
</pre></td></tr></tbody></table></code></pre></figure>

</div>
<div id="canvas-container-4"></div>
</div>

<p>Isn’t that amazing? This is where Schopenhauer and I part ways. Far from being a pessimist, I believe that we are destined for great things. We will will the will (say that three times fast) through our desire to better ourselves, and in doing so, change this negative-sum world into a positive one. If Schopenhauer bumped into you on the street, he would check his pockets, then try to dissuade you from suing. I would apologize and give you a wink.</p>

<p>How would you like to be a Level 5 Mage?</p>

<script src="https://code.jquery.com/jquery-3.7.0.slim.min.js"></script>

<script src="https://cdn.jsdelivr.net/npm/p5@1.1.9/lib/p5.js"></script>

<script src="/assets/blog_scripts/schopenhauer_article.js"></script>

<script> 
    jQuery(".reset-button").click((e) => {
        e.target.innerHTML = "Reset"
        console.log(e.target)
    })
    var i = 0
    var e
    var containers = []
    while((e = document.getElementById("canvas-container-"+i)) && window['container'+i]) {
        var canvas = new p5(window['container'+i], e)
        containers.push(canvas)
        i++
    }
</script>]]></content><author><name>Noel Alemayehu</name></author><category term="schopenhauer" /><category term="programming" /><summary type="html"><![CDATA[How a 19th century crank (possibly) anticipated Java]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="http://noelnegash.github.io/assets/blog_images/post_covers/schopenhauer-cover.jpg" /><media:content medium="image" url="http://noelnegash.github.io/assets/blog_images/post_covers/schopenhauer-cover.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Neural Networks From Scratch With Buzz Lightyear (Part 2: Perceptrons)</title><link href="http://noelnegash.github.io/programming/neural-networks/buzz-lightyear/2020/09/05/buzz-2.html" rel="alternate" type="text/html" title="Neural Networks From Scratch With Buzz Lightyear (Part 2: Perceptrons)" /><published>2020-09-05T18:30:24+00:00</published><updated>2020-09-05T18:30:24+00:00</updated><id>http://noelnegash.github.io/programming/neural-networks/buzz-lightyear/2020/09/05/buzz-2</id><content type="html" xml:base="http://noelnegash.github.io/programming/neural-networks/buzz-lightyear/2020/09/05/buzz-2.html"><![CDATA[<style> .option-button, .reset-button, .change-button {
    border: none;
    color: white;
    text-align: center;
    text-decoration: none;
    display: inline-block;
    font-size: 16px;
    margin: 4px 2px;
    cursor: pointer;
    position: relative;
    font-family: "Montserrat", sans-serif;
    font-size: 14px;
    line-height: 24px;
    font-weight: 700;
    padding: 12px 35px;
    color: #fff;
    text-transform: uppercase;
    border-radius: 3px;
}
.reset-button {
    background-color: #4CAF50; 
    padding: 15px 32px;
}
.change-button {
    background-color: #4274d6;
    display: block;
    padding: 7px 15px;
    }
.option-button {
    background-color: #a12895;
    display: block;
    padding: 7px 15px;
}
.flex-container {
    display: flex;
    flex-direction: row;
    align-items: center;    
    justify-content: center;
}
.flex-container > * {
    margin: 30px;
}
pre {
   text-align: left !important;
}

h4, h3 {
    font-weight: bold;
    margin: 10px auto;
} </style>

<p>In the <a href="/programming/neural-networks/buzz-lightyear/2020/09/05/buzz-1.html">last article</a>, Buzz showed us the basic ideas of calculus and derivatives. In this article, we will use those concepts to make a perceptron.</p>

<p>Before we go ahead and start making one, it’s useful to clarify what a perceptron (and a neural network in general) is. It is an object vaguely inspired by the biological neuron that takes in an input and maps it to a certain output, basically a function. In general, there are two types of perceptrons, ones that output continuous values (called linear regression models) and ones that output discrete values (logistic regression models). Let’s start with the former.</p>

<p>In a linear regression model, we try to optimize a perceptron into making accurate predictions based on a given dataset. The simplest and most common function this perceptron can be is a simple line, ( output = weight cdot input + bias ), weight and bias being the slope and y-intercept. Since we don’t know what weight and bias to use, we set them to random numbers. In the illustration below, you can see what a randomly initialized perceptron acts like and how close it comes to describing the real data, in this case a straight line. The blue lines show the amount of deviation from the correct answer. We call them the <em>cost</em>, <em>loss</em> or <em>error</em> of the perceptron. A convenient method to calculate cost is as the square of the difference between the output and the desired output. It’s convenient because the derivative is easy to calculate.</p>

<div class="flex-container">


<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
11
</pre></td><td class="code"><pre>    <span class="kd">var</span> <span class="nx">perceptron</span> <span class="o">=</span> <span class="p">[</span><span class="nf">random</span><span class="p">(</span><span class="o">-</span><span class="mi">1</span><span class="p">,</span><span class="mi">1</span><span class="p">),</span><span class="nf">random</span><span class="p">(</span><span class="o">-</span><span class="mi">1</span><span class="p">,</span><span class="mi">1</span><span class="p">)]</span>
    <span class="kd">function</span> <span class="nf">calculatePerceptron</span><span class="p">(</span><span class="nx">x</span><span class="p">)</span> <span class="p">{</span>
        <span class="k">return</span> <span class="nx">perceptron</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span><span class="o">*</span><span class="nx">x</span> <span class="o">+</span> <span class="nx">perceptron</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span>
    <span class="p">}</span>
    <span class="kd">function</span> <span class="nf">cost</span><span class="p">(</span><span class="nx">x</span><span class="p">,</span><span class="nx">y</span><span class="p">)</span> <span class="p">{</span>
        <span class="k">return </span><span class="p">(</span><span class="nf">calculatePerceptron</span><span class="p">(</span><span class="nx">x</span><span class="p">)</span> <span class="o">-</span> <span class="nx">y</span><span class="p">)</span><span class="o">**</span><span class="mi">2</span>
    <span class="p">}</span>
    <span class="kd">function</span> <span class="nf">cost_deriv</span><span class="p">(</span><span class="nx">x</span><span class="p">,</span><span class="nx">y</span><span class="p">)</span> <span class="p">{</span>
        
        <span class="k">return</span> <span class="mi">2</span> <span class="o">*</span> <span class="p">(</span><span class="nf">calculatePerceptron</span><span class="p">(</span><span class="nx">x</span><span class="p">)</span> <span class="o">-</span> <span class="nx">y</span><span class="p">)</span>
    <span class="p">}</span>
</pre></td></tr></tbody></table></code></pre></figure>


<div>
    <div id="canvas-container-0"></div>
    <p id="cost0" class="cost"></p>
    <button class="reset-button" onclick="resetPerceptron(0)" type="button"> Reset </button>
</div>

</div>

<p>
In the spirit of this article, let's suppose our data is a map of stars that Buzz Lightyear has to fly to on a mission. He knows the locations of the first few stars in his itinerary, but after that he's on his own. In this scenario, our data is the stars we already know of and the perceptron is meant to fit them as well as predict the location of the unknown stars. Again, since we don't know what weight and bias give us this behavior, we set them to random values. 
</p>

<p>
Since a randomly initialized perceptron is useless, we need to find a way to change it into one with no loss. In other words, we want to find the weight and bias where the loss is at a minimum. If that sounds familiar, that is because it's gradient descent. If we find the derivative of the cost in terms of the weight and bias, we can optimize the perceptron. If the cost is \({((in \cdot w + b) - out)}^2\), the derivative is \(2 \cdot{((in \cdot w + b) - out)}\) in terms of the bias, same as the function above. But for the weight, we need to consider the chain rule. That means we multiply the previous derivative by the derivative of \(in \cdot w \), giving us \(2 \cdot{((in \cdot w + b) - out)} \cdot in\).
</p>

<div class="flex-container">


<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
11
</pre></td><td class="code"><pre><span class="kd">function</span> <span class="nf">update</span><span class="p">()</span> <span class="p">{</span>
    <span class="err">#</span> <span class="nx">variable</span> <span class="nx">to</span> <span class="nx">sum</span> <span class="nx">up</span> <span class="nx">the</span> <span class="nx">derivatives</span>
    <span class="kd">var</span> <span class="nx">adj</span> <span class="o">=</span> <span class="p">[</span><span class="mi">0</span><span class="p">,</span><span class="mi">0</span><span class="p">]</span>
    <span class="k">for </span><span class="p">(</span><span class="kd">var</span> <span class="nx">i</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="nx">i</span> <span class="o">&amp;</span><span class="nx">lt</span><span class="p">;</span> <span class="nx">points</span><span class="p">.</span><span class="nx">length</span><span class="p">;</span> <span class="nx">i</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
        <span class="kd">var</span> <span class="nx">p</span> <span class="o">=</span> <span class="nx">points</span><span class="p">[</span><span class="nx">i</span><span class="p">]</span>
        <span class="nx">adj</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span> <span class="o">-=</span> <span class="nf">cost_deriv</span><span class="p">(</span><span class="nx">p</span><span class="p">.</span><span class="nx">x</span><span class="p">,</span> <span class="nx">p</span><span class="p">.</span><span class="nx">y</span><span class="p">)</span><span class="o">*</span><span class="nx">p</span><span class="p">.</span><span class="nx">x</span>
        <span class="nx">adj</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span> <span class="o">-=</span> <span class="nf">cost_deriv</span><span class="p">(</span><span class="nx">p</span><span class="p">.</span><span class="nx">x</span><span class="p">,</span> <span class="nx">p</span><span class="p">.</span><span class="nx">y</span><span class="p">)</span>
    <span class="p">}</span> 
    <span class="nx">perceptron</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span> <span class="o">+=</span> <span class="mf">0.05</span><span class="o">*</span><span class="nx">adj</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span>
    <span class="nx">perceptron</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span> <span class="o">+=</span> <span class="mf">0.007</span><span class="o">*</span><span class="nx">adj</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span>
<span class="p">}</span>
</pre></td></tr></tbody></table></code></pre></figure>


<div>
    <div id="canvas-container-1"></div>
    <p id="cost1" class="cost"></p>
    <button class="reset-button" onclick="resetPerceptron(1)" type="button"> Start </button>
</div>

</div>

<p>Now that the perceptron has been optimized, it fits the data we have perfectly and and Buzz can find all the other stars he wants by just extending the line.</p>

<p>Note that since the perceptron is a straight line, it can only get a cost of zero when the data itself is in a straight line. If the data has slight deviations, the cost will settle at a number slightly above zero because that is the best possible fit. And if the data isn’t even a line, let’s say it goes up and down, then a single perceptron is basically worthless, since it can only model one direction. You can see the best fits of different types of data on a single perceptron below.</p>

<div class="flex-container">

<div>
    <div id="canvas-container-2"></div>
    <p id="cost2" class="cost"></p>
    <button class="reset-button" onclick="resetPerceptron(2)" type="button"> Start </button>
</div>

<div>
    <button class="change-button" onclick="changeData(2, 'lin')" type="button"> Linear Data </button>
    <button class="change-button" onclick="changeData(2, 'lindev')" type="button"> Linear (Deviation) </button>
    <button class="change-button" onclick="changeData(2, 'quad')" type="button"> Quadratic </button>
    <button class="change-button" onclick="changeData(2, 'quaddev')" type="button"> Quad (Deviation) </button>
    <button class="change-button" onclick="changeData(2, 'cub')" type="button"> Cubic </button>
    <button class="change-button" onclick="changeData(2, 'cubdev')" type="button"> Cubic (Deviation) </button>
    <button class="change-button" onclick="changeData(2, 'sin')" type="button"> Sine </button>
    <button class="change-button" onclick="changeData(2, 'sindev')" type="button"> Sine (Deviation) </button>
</div>

</div>

<p>As you can see, a single perceptron is not expressive enough when it comes to complex data that doesn’t necessarily show linear behavior. In practice, we solve this by using perceptrons in chains and layers to create more complex models, which is what we are going to do in the next and final article. But here, in the spirit of experimentation, we can also change the type of perceptron so that it can fit the data. We can make it a quadratic, cubic or even sine function. And if we get the derivatives right, each one is a valid perceptron that can in some cases outperform a simple linear function.</p>

<div class="flex-container">

<div>
    <div id="canvas-container-3"></div>
    <p id="cost3" class="cost"></p>
    <button class="reset-button" onclick="resetPerceptron(3)" type="button"> Start </button>
</div>

<div>
    <button class="option-button" onclick="containers[3].cp = calculatePerceptron" type="button"> Linear Function </button>
    <button class="option-button" onclick="containers[3].cp = calculatePerceptronQuad" type="button"> Quad Function </button>
    <button class="option-button" onclick="containers[3].cp = calculatePerceptronCub" type="button"> Cubic Function </button>
    <button class="option-button" onclick="containers[3].cp = calculatePerceptronSin" type="button"> Sine Function </button>
</div>

<div>
    <button class="change-button" onclick="changeData(3, 'lin')" type="button"> Linear Data </button>
    <button class="change-button" onclick="changeData(3, 'lindev')" type="button"> Linear (Deviation) </button>
    <button class="change-button" onclick="changeData(3, 'quad')" type="button"> Quadratic </button>
    <button class="change-button" onclick="changeData(3, 'quaddev')" type="button"> Quad (Deviation) </button>
    <button class="change-button" onclick="changeData(3, 'cub')" type="button"> Cubic </button>
    <button class="change-button" onclick="changeData(3, 'cubdev')" type="button"> Cubic (Deviation) </button>
    <button class="change-button" onclick="changeData(3, 'sin')" type="button"> Sine </button>
    <button class="change-button" onclick="changeData(3, 'sindev')" type="button"> Sine (Deviation) </button>
</div>

</div>

<p>We can see a lot of intuitive stuff when we mess around with the function and data pairs. One of the first things we notice is that the cubic function can fit easily to all the types of data. It makes sense, because a quadratic expression is just a cubic expression with a zero coefficient. And a linear expression is one with two. If that’s the case, you might be wondering why we don’t prefer them to linear expressions in neural networks. The correct answer is that they tend to overfit and have exploding gradients, both of which will be explained in the next article. The simplified answer is that cubic functions, and higher-polynomials in general, have very sensitive coefficients. A small change affects the graph drastically, making them too unstable for practical purposes. And even when we find a way to make them stable, it comes at the cost of speed and efficiency, which have a lot of importance when it comes to practical machine learning, since it’s not uncommon for a neural network to have thousands, millions or billions of perceptrons.</p>

<p>Another less obvious thing we notice is that the sine function sometimes has trouble fitting to quadratic data, remaining a horizontal line. This is because the perceptron has found a <em>local</em> minimum instead of the <em>absolute</em> minimum. What it means is that sometimes your perceptron finds the best weights if you initialize it a certain way and misses them in others, which also comes into play when you make large networks.</p>

<p>All of this applies to logistic regression as well. But while linear regression makes perceptrons predict data, logistic regression makes them classify it. For example, if the galaxy is either being invaded or experiencing a pandemic, Buzz doesn’t need to know where the spaceships or planets are, he needs to know <em>what</em> they are when he sees them. In this scenario, we have a dataset of green(safe) and orange(dangerous) points, the few places we know the status of, and Buzz wants to figure out which group a newly added point belongs to.</p>

<p>
So how do we make a perceptron output discrete values? We can't. But we can do the next best thing, we can make it output probabilities on those discrete values. Linear perceptrons are not very useful to decide probabilities, because it makes very little sense to say 1000% green or -12% orange. We need it to output something from 0% to 100%. This is where the sigmoid activation function comes in. An activation function is a sort of wrapper we put over the perceptron to make its output convenient to use. Sigmoid uses the formula \( \dfrac{1}{1+e^{-x}} \) to squash a linear equation between 0 and 1. Now 1000 maps to 100% and -12 maps to 0%. Sigmoid also has a convenient derivative, \( sigmoid(x) \cdot (1-sigmoid(x)) \).
</p>

<div class="flex-container">

<div id="canvas-container-4"></div>
</div>

<p>
When we use logistic regression, it's common to change the cost function too. While we can technically still use the difference squared, it gives poor results because the maximum possible cost is only 1, making gradient descent work slowly. We use the negative log functions, \(-log({1 - output})\) and \(-log({output})\), depending on whether we want 0 or 1, to give us a more exaggerated cost, as you can see below.
</p>

<div class="flex-container">

<div id="canvas-container-5"></div>
</div>

<div class="flex-container">

<div id="canvas-container-6"></div>
</div>

<p>
Combining the activation function and cost function with a simple linear perceptron gives us a more complex derivative because of all the nesting. But if we follow the chain rule, it comes together easily enough. The cost function is \(log(sig(perceptron))\), so the derivative will be \(logDeriv(sig(perceptron)) \cdot sigDeriv(perceptron) \cdot perceptronDeriv \). Here it is in code form, finding a line to separate the green and orange points.
</p>

<div class="flex-container">

<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
</pre></td><td class="code"><pre><span class="kd">function</span> <span class="nf">sigmoid</span><span class="p">(</span><span class="nx">x</span><span class="p">)</span> <span class="p">{</span>
    <span class="k">return</span> <span class="mi">1</span><span class="o">/</span><span class="p">(</span><span class="mi">1</span><span class="o">+</span><span class="nb">Math</span><span class="p">.</span><span class="nf">exp</span><span class="p">(</span><span class="o">-</span><span class="nx">x</span><span class="p">))</span>
<span class="p">}</span>
<span class="kd">function</span> <span class="nf">sigmoid_deriv</span><span class="p">(</span><span class="nx">x</span><span class="p">)</span> <span class="p">{</span>
    <span class="k">return</span> <span class="nf">sigmoid</span><span class="p">(</span><span class="nx">x</span><span class="p">)</span><span class="o">*</span><span class="p">(</span><span class="mi">1</span><span class="o">-</span><span class="nf">sigmoid</span><span class="p">(</span><span class="nx">x</span><span class="p">))</span>
<span class="p">}</span>
<span class="kd">function</span> <span class="nf">calculatePerceptron</span><span class="p">(</span><span class="nx">data</span><span class="p">)</span> <span class="p">{</span>
    <span class="k">return</span> <span class="nx">perceptron</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span><span class="o">*</span><span class="nx">data</span><span class="p">.</span><span class="nx">x</span> <span class="o">+</span> <span class="nx">perceptron</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span><span class="o">*</span><span class="nx">data</span><span class="p">.</span><span class="nx">y</span> <span class="o">+</span> <span class="nx">perceptron</span><span class="p">[</span><span class="mi">2</span><span class="p">])</span>
<span class="p">}</span>
<span class="kd">function</span> <span class="nf">cost</span><span class="p">(</span><span class="nx">data</span><span class="p">)</span> <span class="p">{</span>
    <span class="kd">var</span> <span class="nx">x</span> <span class="o">=</span> <span class="nf">sigmoid</span><span class="p">(</span><span class="nf">calculatePerceptron</span><span class="p">(</span><span class="nx">data</span><span class="p">))</span>

    <span class="k">if </span><span class="p">(</span><span class="nx">data</span><span class="p">.</span><span class="kd">class</span> <span class="err">== 1) {
            </span><span class="nc">return</span> <span class="o">-</span><span class="nb">Math</span><span class="p">.</span><span class="nf">log</span><span class="p">(</span><span class="nx">x</span><span class="p">)</span>
    <span class="p">}</span>
    <span class="k">return</span> <span class="o">-</span><span class="nb">Math</span><span class="p">.</span><span class="nf">log</span><span class="p">(</span><span class="mi">1</span><span class="o">-</span><span class="nx">x</span><span class="p">)</span>
<span class="p">}</span>
<span class="kd">function</span> <span class="nf">cost_deriv</span><span class="p">(</span><span class="nx">data</span><span class="p">,</span> <span class="nx">y</span><span class="p">)</span> <span class="p">{</span>
    <span class="kd">var</span> <span class="nx">x</span> <span class="o">=</span> <span class="nf">sigmoid</span><span class="p">(</span><span class="nf">calculatePerceptron</span><span class="p">(</span><span class="nx">data</span><span class="p">))</span>

    <span class="k">if </span><span class="p">(</span><span class="nx">data</span><span class="p">.</span><span class="kd">class</span> <span class="err">== 1) {
            </span><span class="nc">return</span> <span class="o">-</span><span class="mi">1</span><span class="o">/</span><span class="nx">x</span>
    <span class="p">}</span>
    <span class="k">return</span> <span class="mi">1</span><span class="o">/</span><span class="p">(</span><span class="mi">1</span><span class="o">-</span><span class="nx">x</span><span class="p">)</span> 
    <span class="c1">// not negative because the inner derivative cancels it out</span>
<span class="p">}</span>
<span class="kd">function</span> <span class="nf">update</span><span class="p">()</span> <span class="p">{</span>
    <span class="c1">// variable to sum up the derivatives</span>
    <span class="kd">var</span> <span class="nx">adj</span> <span class="o">=</span> <span class="p">[</span><span class="mi">0</span><span class="p">,</span><span class="mi">0</span><span class="p">,</span><span class="mi">0</span><span class="p">]</span>
    <span class="k">for </span><span class="p">(</span><span class="kd">var</span> <span class="nx">i</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="nx">i</span> <span class="o">&amp;</span><span class="nx">lt</span><span class="p">;</span> <span class="nx">points</span><span class="p">.</span><span class="nx">length</span><span class="p">;</span> <span class="nx">i</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
        <span class="kd">var</span> <span class="nx">p</span> <span class="o">=</span> <span class="nx">points</span><span class="p">[</span><span class="nx">i</span><span class="p">]</span>
        <span class="nx">adj</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span> <span class="o">-=</span> <span class="nf">cost_deriv</span><span class="p">(</span><span class="nx">p</span><span class="p">)</span> <span class="o">*</span> <span class="nf">sigmoid_deriv</span><span class="p">(</span><span class="nf">calculatePerceptron</span><span class="p">(</span><span class="nx">p</span><span class="p">))</span> <span class="o">*</span> <span class="nx">p</span><span class="p">.</span><span class="nx">x</span>
        <span class="nx">adj</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span> <span class="o">-=</span> <span class="nf">cost_deriv</span><span class="p">(</span><span class="nx">p</span><span class="p">)</span> <span class="o">*</span> <span class="nf">sigmoid_deriv</span><span class="p">(</span><span class="nf">calculatePerceptron</span><span class="p">(</span><span class="nx">p</span><span class="p">))</span> <span class="o">*</span> <span class="nx">p</span><span class="p">.</span><span class="nx">y</span>
        <span class="nx">adj</span><span class="p">[</span><span class="mi">2</span><span class="p">]</span> <span class="o">-=</span> <span class="nf">cost_deriv</span><span class="p">(</span><span class="nx">p</span><span class="p">)</span> <span class="o">*</span> <span class="nf">sigmoid_deriv</span><span class="p">(</span><span class="nf">calculatePerceptron</span><span class="p">(</span><span class="nx">p</span><span class="p">))</span>
    <span class="p">}</span> 
    <span class="nx">perceptron</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span> <span class="o">+=</span> <span class="mf">0.3</span><span class="o">*</span><span class="nx">adj</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span>
    <span class="nx">perceptron</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span> <span class="o">+=</span> <span class="mf">0.3</span><span class="o">*</span><span class="nx">adj</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span>
    <span class="nx">perceptron</span><span class="p">[</span><span class="mi">2</span><span class="p">]</span> <span class="o">+=</span> <span class="mf">0.3</span><span class="o">*</span><span class="nx">adj</span><span class="p">[</span><span class="mi">2</span><span class="p">]</span>
<span class="p">}</span>
</pre></td></tr></tbody></table></code></pre></figure>

<div>
    <div id="canvas-container-7"></div>
    <p id="cost7" class="cost"></p>
    <button class="reset-button" onclick="resetPerceptron(7)" type="button"> Start </button>
</div>
</div>

<p>Now that we have the ideal perceptron, all we need to do to find the class of a data point is to check the ouptut of the sigmoid. If it’s 0.5, then the point lies right on the line and can’t be said to be one or the other. If it is less than 0.5, we say it is closer to the 0 class and if it is greater we say it is in the 1 class. We can visualize this by trying the perceptron on the whole plane and seeing how the output changes.</p>

<div class="flex-container">
<div>
    <div id="canvas-container-8"></div>
    <p id="cost8" class="cost"></p>
    <button class="reset-button" onclick="resetPerceptron(8)" type="button"> Start </button>
</div>
</div>

<p>Now Buzz can tell whether any section of space is safe or dangerous based on the little data he has. Let’s see how it works on different datasets. You can change the datasets below to random clusters and nested rings to see how well the perceptron manages to separate them.</p>

<div class="flex-container">
<div>
    <div id="canvas-container-9"></div>
    <p id="cost9" class="cost"></p>
    <button class="reset-button" onclick="resetPerceptron(9)" type="button"> Start</button>
</div>

<div>
    <button class="change-button" onclick="changeData(9, 'clust')" type="button"> Two Clusters </button>
    <button class="change-button" onclick="changeData(9, 'clustin')" type="button"> Two Clusters (Intersecting) </button>
    <button class="change-button" onclick="changeData(9, '3clust')" type="button"> Three Clusters </button>
    <button class="change-button" onclick="changeData(9, '3clustin')" type="button"> Three Clusters (Intersecting) </button>
    <button class="change-button" onclick="changeData(9, 'rings')" type="button"> Two Rings </button>
    <button class="change-button" onclick="changeData(9, '3rings')" type="button"> Three Rings </button>
</div>
</div>

<p>Like the linear regression example, sometimes a simple perceptron isn’t complex enough to model data with curves, especially closed curves like circles, which can never be modeled by a line. The cost never reaches a reasonable low either. Again, in practice this is remedied by using multiple layers of simple perceptrons, but for the sake of experimentation, we will see how more exotic perceptrons fare with the same data.</p>

<div class="flex-container">
<div>
    <div id="canvas-container-10"></div>
    <p id="cost10" class="cost"></p>
    <button class="reset-button" onclick="resetPerceptron(10)" type="button"> Start</button>
</div>

<div>
    <button class="option-button" onclick="containers[10].cp = calculatePerceptronSigmoid" type="button"> Linear Function </button>
    <button class="option-button" onclick="containers[10].cp = calculatePerceptronQuadSigmoid" type="button"> Quad Function </button>
    <button class="option-button" onclick="containers[10].cp = calculatePerceptronCubSigmoid" type="button"> Cubic Function </button>
    <button class="option-button" onclick="containers[10].cp = calculatePerceptronSinSigmoid" type="button"> Sine Function </button>
    <button class="option-button" onclick="containers[10].cp = calculatePerceptronSinQuadSigmoid" type="button"> Sine(Quad) Function </button>
</div>

<div>
    <button class="change-button" onclick="changeData(10, 'clust')" type="button"> Two Clusters </button>
    <button class="change-button" onclick="changeData(10, 'clustin')" type="button"> Two Clusters (Intersecting) </button>
    <button class="change-button" onclick="changeData(10, '3clust')" type="button"> Three Clusters </button>
    <button class="change-button" onclick="changeData(10, '3clustin')" type="button"> Three Clusters (Intersecting) </button>
    <button class="change-button" onclick="changeData(10, 'rings')" type="button"> Two Rings </button>
    <button class="change-button" onclick="changeData(10, '3rings')" type="button"> Three Rings </button>
</div>
</div>

<p>As you can see, more complex functions tend to fit more complex data, but even the most complex functions can’t perfectly separate data that is thoroughly mixed. This is a common problem you will see with neural networks that are given bad or patternless data, they will give back confusing information that contains a lot of false positives and negatives.</p>

<p>
If you are wondering why the quadratic function can model circles and ellipses, it is because we are using the the general formula \( a \cdot x^2 + b \cdot x+c \cdot y^2+d \cdot y + e \), which can model those shapes as well as parabolic and hyperbolic curves. Plugging it into a sine function adds another layer of complexity, resulting in graphs like concentric rings and expanding stars.
</p>

<p>And that is it for perceptrons. Read <a href="/programming/neural-networks/buzz-lightyear/2020/09/05/buzz-3.html">the next article</a> to see how we can combine perceptrons to create much more powerful and expressive neural networks.</p>

<script src="https://code.jquery.com/jquery-3.7.0.slim.min.js"></script>

<script src="https://cdn.jsdelivr.net/npm/p5@1.1.9/lib/p5.js"></script>

<script src="https://cdn.mathjax.org/mathjax/latest/MathJax.js?config=TeX-AMS-MML_HTMLorMML" type="text/javascript"></script>

<script src="/assets/blog_scripts/perceptron_tutorial.js"></script>

<script> 
    jQuery(".reset-button").click((e) => {
            e.target.innerHTML = "Reset"
        console.log(e.target)
    })
    var i = 0
    var e
    var containers = []
    while((e = document.getElementById("canvas-container-"+i)) && window['container'+i]) {
        var canvas = new p5(window['container'+i], e)
        containers.push(canvas)
        i++
    }
</script>]]></content><author><name>Noel Alemayehu</name></author><category term="programming" /><category term="neural-networks" /><category term="buzz-lightyear" /><summary type="html"><![CDATA[Buzz Lightyear teaches math]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="http://noelnegash.github.io/assets/blog_images/post_covers/buzz-cover.png" /><media:content medium="image" url="http://noelnegash.github.io/assets/blog_images/post_covers/buzz-cover.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Neural Networks From Scratch With Buzz Lightyear (Part 3: Multi-Layer Perceptrons)</title><link href="http://noelnegash.github.io/programming/neural-networks/buzz-lightyear/2020/09/05/buzz-3.html" rel="alternate" type="text/html" title="Neural Networks From Scratch With Buzz Lightyear (Part 3: Multi-Layer Perceptrons)" /><published>2020-09-05T18:30:24+00:00</published><updated>2020-09-05T18:30:24+00:00</updated><id>http://noelnegash.github.io/programming/neural-networks/buzz-lightyear/2020/09/05/buzz-3</id><content type="html" xml:base="http://noelnegash.github.io/programming/neural-networks/buzz-lightyear/2020/09/05/buzz-3.html"><![CDATA[<style> .option-button, .reset-button, .change-button {
    border: none;
    color: white;
    text-align: center;
    text-decoration: none;
    display: inline-block;
    font-size: 16px;
    margin: 4px 2px;
    cursor: pointer;
    position: relative;
    font-family: "Montserrat", sans-serif;
    font-size: 14px;
    line-height: 24px;
    font-weight: 700;
    padding: 12px 35px;
    color: #fff;
    text-transform: uppercase;
    border-radius: 3px;
}
.reset-button {
    background-color: #4CAF50; 
    padding: 15px 32px;
}
.change-button {
    background-color: #4274d6;
    display: block;
    padding: 7px 15px;
    }
.option-button {
    background-color: #a12895;
    display: block;
    padding: 7px 15px;
}
.flex-container {
    display: flex;
    flex-direction: row;
    align-items: center;    
    justify-content: center;
}
.flex-container > * {
    margin: 30px;
}
pre {
   text-align: left !important;
}

h4, h3 {
    font-weight: bold;
    margin: 10px auto;
} </style>

<p>In the <a href="/programming/neural-networks/buzz-lightyear/2020/09/05/buzz-2.html">previous article</a>, we learned how to make a single perceptron find a best-fit line, with an introduction to complex models and activation functions. In this article, we will expand those concepts into creating a fully-connected multi-layer neural network by chaining individual perceptrons. This is accomplished by combining gradient descent with backpropagation.</p>

<p>A multi-layer perceptron(MLP) is composed of layers, each of which contains a certain amount of perceptrons. When we input some data, each layer calculates an output based on its perceptrons’ weights and values, but instead of passing those outputs as the final answer, it passes them onto the next layer for more computation until it reaches the final layer. We call this computation of an output a <em>forward pass</em>. The benefit of all this convolution is that it has lots of connections, which means a lot of complexity. As we have already seen, complexity is very useful in a neural network, because it allows us to fit any type of data.</p>

<p>In the simulation below, you can see a visual representation of an MLP. The buttons allow you to choose the number of layers and perceptrons per layer. You can also press any perceptron in the simulation to see how the forward pass goes from there.</p>

<div class="flex-container">
<div id="canvas-container-0"></div>

<div>
    <button class="option-button" onclick="containers[0].layers(3)" type="button"> 3 Layers </button>
    <button class="option-button" onclick="containers[0].layers(4)" type="button"> 4 Layers </button>
    <button class="option-button" onclick="containers[0].layers(5)" type="button"> 5 Layers </button>
</div>

<div>
    <button class="change-button" onclick="containers[0].perceptrons(1)" type="button"> 1 Perceptron </button>
    <button class="change-button" onclick="containers[0].perceptrons(2)" type="button"> 2 Perceptrons </button>
    <button class="change-button" onclick="containers[0].perceptrons(3)" type="button"> 3 Perceptrons </button>
    <button class="change-button" onclick="containers[0].perceptrons(4)" type="button"> 4 Perceptrons </button>
</div>

</div>

<p>Note that the number of perceptrons can change from layer to layer. In fact, some very useful variants of neural network such as autoencoders and GANs depend on this property.</p>

<div class="flex-container">
<div id="canvas-container-1"></div>
<div>
    <button class="option-button" onclick="containers[1].funnel()" type="button"> Funnel Shape </button>
    <button class="option-button" onclick="containers[1].diamond()" type="button"> Diamond Shape </button>
    <button class="option-button" onclick="containers[1].autoencoder()" type="button"> Autoencoder </button>
</div>
</div>

<p>Because there are so many weights, biases, inputs and biases involved in a neural network, it becomes practically impossible to manipulate them as individual values. This is why machine learning libraries (and us, starting now) use arrays, more accurately matrices and tensors, to represent an MLP. For simplicity’s sake, every matrix used in this article will be one dimensional. It is also important to note that all matrix multiplication in the code will be element-wise, meaning</p>

<div class="flex-container">
<p>
\(\begin{pmatrix} x_1 \\ x_2 \end{pmatrix} \cdot \begin{pmatrix} y_1 \\ y_2 \end{pmatrix} = \begin{pmatrix} x_1 \cdot y_1 \\ x_2 \cdot y_2 \end{pmatrix}\)
</p>
</div>

<p>Each perceptron will be a matrix of weights, one for each input (output of the previous layer), and one bias. That way, we can find the output of a perceptron by multiplying in with a matrix of the input data and adding the bias, like so.</p>

<div class="flex-container">
<p>
\(\sum({\begin{pmatrix} w_1 \\ w_2 \end{pmatrix} \cdot \begin{pmatrix} x_1 \\ x_2 \end{pmatrix}}) + b \\ \sum({\begin{pmatrix} w_1 \cdot x_1 \\ w_2 \cdot x_2 \end{pmatrix}}) + b \\ w_1 \cdot x_1 + w_2 \cdot x_2 + b\)
</p>
</div>

<p>In the same vein, we can define a layer as an array of perceptrons and an MLP as an array of layers. By describing them in this way, we significantly reduce the number of variables we have to work with and simplify our manipulations on them (see the <code class="language-plaintext highlighter-rouge">eval</code> functions in the accompanying code).  Now we can calculate perceptrons and layers with thousands of variables in just one function call. In the simulation below, you can hover over each perceptron below to see how many parameters(weights+bias) it has.</p>

<div class="flex-container">

<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
</pre></td><td class="code"><pre>    <span class="c1">// a percptron is an array of weights</span>
    <span class="kd">class</span> <span class="nc">Perceptron</span> <span class="kd">extends</span> <span class="nc">Array</span><span class="p">{</span>
        <span class="nf">constructor </span><span class="p">(</span><span class="nx">numInputs</span><span class="p">)</span> <span class="p">{</span>
            <span class="k">for </span><span class="p">(</span><span class="kd">var</span> <span class="nx">i</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="nx">i</span> <span class="o">&amp;</span><span class="nx">lt</span><span class="p">;</span> <span class="nx">numInputs</span><span class="p">;</span> <span class="nx">i</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
                <span class="k">this</span><span class="p">.</span><span class="nf">push</span><span class="p">(</span><span class="nf">random</span><span class="p">(</span><span class="o">-</span><span class="mi">1</span><span class="p">,</span><span class="mi">1</span><span class="p">))</span>
            <span class="p">}</span>
            <span class="k">this</span><span class="p">.</span><span class="nx">bias</span> <span class="o">=</span> <span class="nf">random</span><span class="p">(</span><span class="o">-</span><span class="mi">1</span><span class="p">,</span><span class="mi">1</span><span class="p">)</span>
        <span class="p">}</span>
        <span class="nb">eval</span> <span class="o">=</span> <span class="nf">function </span><span class="p">(</span><span class="nx">d</span><span class="p">)</span> <span class="p">{</span>
            <span class="k">return</span> <span class="nf">sum</span><span class="p">(</span><span class="k">this</span> <span class="o">*</span> <span class="nx">d</span><span class="p">)</span> <span class="o">+</span> <span class="k">this</span><span class="p">.</span><span class="nx">bias</span>
        <span class="p">}</span>
    <span class="p">}</span>

    <span class="c1">// likewise, a layer is an array of perceptrons</span>
    <span class="kd">class</span> <span class="nc">Layer</span> <span class="kd">extends</span> <span class="nc">Array</span> <span class="p">{</span>
        <span class="nf">constructor </span><span class="p">(</span><span class="nx">numPerceptrons</span><span class="p">,</span> <span class="nx">numInputs</span><span class="p">)</span> <span class="p">{</span>
            <span class="k">super</span><span class="p">()</span>
            <span class="k">for </span><span class="p">(</span><span class="kd">var</span> <span class="nx">i</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="nx">i</span> <span class="o">&amp;</span><span class="nx">lt</span><span class="p">;</span> <span class="nx">numPerceptrons</span><span class="p">;</span> <span class="nx">i</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
                <span class="k">this</span><span class="p">.</span><span class="nf">push</span><span class="p">(</span><span class="k">new</span> <span class="nc">Perceptron</span><span class="p">(</span><span class="nx">numInputs</span><span class="p">))</span>
            <span class="p">}</span>
        <span class="p">}</span>
        <span class="nb">eval</span> <span class="o">=</span> <span class="nf">function </span><span class="p">(</span><span class="nx">d</span><span class="p">)</span> <span class="p">{</span>
            <span class="kd">var</span> <span class="nx">res</span> <span class="o">=</span> <span class="p">[]</span>
            <span class="k">this</span><span class="p">.</span><span class="nx">lastInput</span> <span class="o">=</span> <span class="nx">d</span>
            <span class="k">for </span><span class="p">(</span><span class="kd">var</span> <span class="nx">i</span> <span class="o">&amp;</span><span class="nx">lt</span><span class="p">;</span> <span class="mi">0</span><span class="p">;</span> <span class="nx">i</span> <span class="o">&amp;</span><span class="nx">lt</span><span class="p">;</span> <span class="k">this</span><span class="p">.</span><span class="nx">length</span><span class="p">;</span> <span class="nx">i</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
                <span class="nx">res</span><span class="p">.</span><span class="nf">push</span><span class="p">(</span><span class="k">this</span><span class="p">[</span><span class="nx">i</span><span class="p">].</span><span class="nf">eval</span><span class="p">(</span><span class="nx">d</span><span class="p">))</span>
            <span class="p">}</span>
            <span class="k">this</span><span class="p">.</span><span class="nx">lastOutput</span> <span class="o">=</span> <span class="nx">d</span>
            <span class="k">return</span> <span class="nx">res</span>
        <span class="p">}</span>
    <span class="p">}</span>

    <span class="c1">// finally, an MLP is an array of layers</span>
    <span class="kd">class</span> <span class="nc">MLP</span> <span class="kd">extends</span> <span class="nc">Array</span> <span class="p">{</span>
        <span class="nf">constructor </span><span class="p">(</span><span class="nx">schema</span><span class="p">)</span> <span class="p">{</span>
            <span class="k">super</span><span class="p">()</span>
            <span class="k">for </span><span class="p">(</span><span class="kd">var</span> <span class="nx">i</span> <span class="o">=</span> <span class="mi">1</span><span class="p">;</span> <span class="nx">i</span> <span class="o">&amp;</span><span class="nx">lt</span><span class="p">;</span> <span class="nx">schema</span><span class="p">.</span><span class="nx">length</span><span class="p">;</span> <span class="nx">i</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
                <span class="k">this</span><span class="p">.</span><span class="nf">push</span><span class="p">(</span><span class="k">new</span> <span class="nc">Layer</span><span class="p">(</span><span class="nx">schema</span><span class="p">[</span><span class="nx">i</span><span class="p">],</span> <span class="nx">schema</span><span class="p">[</span><span class="nx">i</span><span class="o">-</span><span class="mi">1</span><span class="p">]))</span>
            <span class="p">}</span>
        <span class="p">}</span>
    <span class="p">}</span>

    <span class="c1">// first element describes the incoming data</span>
    <span class="c1">// the rest describe the number of perceptrons per layer</span>
    <span class="kd">var</span> <span class="nx">schema</span> <span class="o">=</span> <span class="p">[</span><span class="mi">1</span><span class="p">,</span><span class="mi">1</span><span class="p">,</span><span class="mi">2</span><span class="p">,</span><span class="mi">3</span><span class="p">,</span><span class="mi">4</span><span class="p">]</span>
    <span class="kd">var</span> <span class="nx">mlp</span> <span class="o">=</span> <span class="k">new</span> <span class="nc">MLP</span><span class="p">(</span><span class="nx">schema</span><span class="p">)</span>
</pre></td></tr></tbody></table></code></pre></figure>

<div id="canvas-container-2"></div>
</div>

<p>Since we have put down our MLP components in code, we have all we need to implement the forward pass algorithm and see how the it maps out on a graph.</p>

<div class="flex-container">

<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
11
12
</pre></td><td class="code"><pre>    <span class="nx">MLP</span><span class="p">.</span><span class="nx">forward</span> <span class="o">=</span> <span class="nf">function </span><span class="p">(</span><span class="nx">input</span><span class="p">)</span> <span class="p">{</span>

        <span class="c1">// feed forward the output of current layer</span>
        <span class="c1">// as input of next layer</span>
        <span class="k">for </span><span class="p">(</span><span class="kd">var</span> <span class="nx">i</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="nx">i</span> <span class="o">&amp;</span><span class="nx">lt</span><span class="p">;</span> <span class="k">this</span><span class="p">.</span><span class="nx">length</span><span class="p">;</span> <span class="nx">i</span><span class="o">++</span><span class="p">)</span>
            <span class="nx">input</span> <span class="o">=</span> <span class="k">this</span><span class="p">[</span><span class="nx">i</span><span class="p">].</span><span class="nf">eval</span><span class="p">(</span><span class="nx">input</span><span class="p">)</span>

        <span class="k">return</span> <span class="nx">input</span>
    <span class="p">}</span>
    
    <span class="kd">var</span> <span class="nx">schema</span> <span class="o">=</span> <span class="p">[</span><span class="mi">1</span><span class="p">,</span><span class="mi">1</span><span class="p">,</span><span class="mi">2</span><span class="p">,</span><span class="mi">3</span><span class="p">,</span><span class="mi">4</span><span class="p">]</span>
    <span class="kd">var</span> <span class="nx">mlp</span> <span class="o">=</span> <span class="k">new</span> <span class="nc">MLP</span><span class="p">(</span><span class="nx">schema</span><span class="p">)</span>
</pre></td></tr></tbody></table></code></pre></figure>

<div>
    <div id="canvas-container-3"></div>
    <button class="reset-button" onclick="currentContainer=3;containers[3].mlp = new MLP([1,1,2,3,4,1])"> Reset </button>
</div>

</div>

<p>You might be thinking “Huh? Isn’t the whole point of an MLP that it’s more complex than a line? What’s the difference between this and a single perceptron?”. That is 100% correct. Nesting perceptrons on their own is just applying transformations on the line, so it doesn’t change the shape of the graph. This is where ReLU comes in. ReLU is an activation function (like sigmoid) that takes a line and clips off its negative side. You can see what appling ReLU does to random lines below.</p>

<div class="flex-container">
<div>
    <div id="canvas-container-4"></div>
    <button class="reset-button" onclick="containers[4].m = containers[4].random(-1,1);containers[4].b = containers[4].random(-1,1);"> Reset </button>
</div>
</div>

<p>While it might seem like a trivial change, it is enough to make our MLP adopt shapes much more complex than quadratic and cubic functions, because each perceptron has potential to change the direction of the line. The derivative of a ReLU is also very convenient. It is 1 for the line segment, and 0 for the flattened segment.</p>

<div class="flex-container">

<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
4
</pre></td><td class="code"><pre>    <span class="c1">// we don't make the last layer a ReLU </span>
    <span class="c1">// because we want the output to allow negative values</span>
    <span class="kd">var</span> <span class="nx">schema</span> <span class="o">=</span> <span class="p">[</span><span class="mi">1</span><span class="p">,</span> <span class="p">[</span><span class="mi">1</span><span class="p">,</span><span class="dl">'</span><span class="s1">relu</span><span class="dl">'</span><span class="p">],</span> <span class="p">[</span><span class="mi">2</span><span class="p">,</span><span class="dl">'</span><span class="s1">relu</span><span class="dl">'</span><span class="p">],</span> <span class="p">[</span><span class="mi">3</span><span class="p">,</span><span class="dl">'</span><span class="s1">relu</span><span class="dl">'</span><span class="p">],</span> <span class="mi">4</span><span class="p">]</span>
    <span class="kd">var</span> <span class="nx">mlp</span> <span class="o">=</span> <span class="k">new</span> <span class="nc">MLP</span><span class="p">(</span><span class="nx">schema</span><span class="p">)</span>
</pre></td></tr></tbody></table></code></pre></figure>

<div>
    <div id="canvas-container-5"></div>
    <button class="reset-button" onclick="currentContainer=5;containers[5].mlp = new MLP([1,[1,'relu'],[2,'relu'],[3,'relu'],4,1])"> Reset </button>
</div>
</div>

<p>
Speaking of derivatives, I think this is a good place to refresh and go a little deeper into the calculus we have glossed over. If you already know calculus, feel free to skip this section. If you don't, let's go back and talk about how derivatives are the slope of a curve at a certain point. Since they are slopes, we can represent them as fractions. Like we say \( m = \dfrac{\Delta y}{\Delta x} \) for a line, we can say  \( m = \dfrac{\partial y}{\partial x} \) for any graph in general (lines included). It is read as the derivative of "y in terms of x". In a normal perceptron, we know that our cost function is \( c(x) = {(x-y)}^2 \), its derivative in terms of \(x\) is \(2 \cdot (x-y) \) and our perceptron is \(p(x) = w \cdot input + b\). In a normal perceptron, the cost would be \(c(p(x))\) and we are supposed to find the derivative of \(c(p(x))\) in terms of w (i.e \(\dfrac{\partial c(p(x))}{\partial w}\)).
</p>

<p>
Since we only know it in terms of \(p(x)\) (i.e \(\dfrac{\partial c(p(x))}{\partial p(x)} = 2 \cdot (p(x)-y)\)), it looks like we are at a dead end. But this is where a convenient property of fractions comes into play. We can use the property \(\dfrac{y}{x}\cdot\dfrac{x}{z} = \dfrac{y}{z}\) to rephrase \(\dfrac{\partial c(p(x))}{\partial w}\) as \(\dfrac{\partial c(p(x))}{\partial p(x)} \cdot \dfrac{\partial p(x)}{\partial w}\) which simplifies into \(2 \cdot (p(x)-y) \cdot input \). Seem familiar? It's the chain rule. I showed it to you in this slightly more verbose way to explain how backpropagation works.
</p>

<p>
Backpropagation is an algorithm that uses the chan rule of calculus to calculate derivatives for perceptrons within perceptrons or, more intuitively, perceptrons that feed into other perceptrons. Let's say you have two perceptrons(\(p_1\) and \(p_2\)) joined end to end. That gives you a structure and equations that look like this.
</p>

<div class="flex-container">
    <p style="height:400px; padding:170px 0;">
    \(p_1 = w_1 \cdot x + b_1 \\ p_2 = w_2 \cdot p_1 + b_2 \\ cost = (p_2 - y)^2\)</p>
    <div id="canvas-container-6"></div>
</div>

<p>
To do gradient descent, we need to be able to find the derivative of the cost with respect to both perceptrons. Let's start with the easier \(p_2\).
</p>

<div class="flex-container">
<p>\(\dfrac{\partial cost}{\partial p_2} = 2 \cdot (p_2 - y)\)</p>
</div>

<p>
How about \(p_1\)? We can use chain rule like so.
</p>

<div class="flex-container">
<p>\(\dfrac{\partial cost}{\partial p_1} = \dfrac{\partial cost}{\partial p_2} \cdot \dfrac{\partial p_2}{\partial p_1} = 2 \cdot (p_2 - y) \cdot w_2\) </p>
</div>

<p>See the pattern? The derivative for one perceptron is depends on the one immediately next to it. Extending the chain to three perceptrons makes it even clearer.</p>

<div class="flex-container">
    <p style="height:400px; padding:170px 0;">\(p_1 = w_1 \cdot x + b_1 \\ p_2 = w_2 \cdot p_1 + b_2 \\ p_3 = w_3 \cdot p_2 + b_3\\ cost = (p_3 - y)^2\) </p>
    <div id="canvas-container-7"></div>
</div>

<div class="flex-container">
<p>\(\dfrac{\partial cost}{\partial p_3} = 2 \cdot (p_3 - y)\) </p>
</div>

<div class="flex-container">
<p>\(\dfrac{\partial cost}{\partial p_2} = \dfrac{\partial cost}{\partial p_3} \cdot \dfrac{\partial p_3}{\partial p_2} = 2 \cdot (p_3 - y) \cdot w_3\) </p>
</div>

<div class="flex-container">
<p>\(\dfrac{\partial cost}{\partial p_1} = \dfrac{\partial cost}{\partial p_3} \cdot \dfrac{\partial p_3}{\partial p_2}\cdot \dfrac{\partial p_2}{\partial p_1} =2 \cdot (p_3 - y) \cdot w_3 \cdot w_2\)</p>
</div>

<p>Let’s analyze one last structure before we put this into code. What happens when a perceptron feeds into two instead of one?</p>

<div class="flex-container">
<p style="height:400px; padding:170px 0;">\(p_1 = w_1 \cdot x + b_1 \\ p_2 = w_2 \cdot p_1 + b_2 \\ p_3 = w_3 \cdot p_1 + b_3 \\ p_4 = w_4 \cdot p_2 + w_5 \cdot p_3 + b \\ cost = (p_4 - y)^2\) </p>
<div id="canvas-container-8"></div>
</div>

<div class="flex-container">
<p>\(\dfrac{\partial cost}{\partial p_4} = 2 \cdot (p_4 - y)\) </p>
</div>

<div class="flex-container">
<p>\(\dfrac{\partial cost}{\partial p_2} = \dfrac{\partial cost}{\partial p_4} \cdot \dfrac{\partial p_4}{\partial p_2} = 2 \cdot (p_4 - y) \cdot w_4 \ \dfrac{\partial cost}{\partial p_3} = \dfrac{\partial cost}{\partial p_4} \cdot \dfrac{\partial p_4}{\partial p_3} = 2 \cdot (p_4 - y) \cdot w_5\)</p>
</div>

<p>Since P1 has both a presence in P2 and P3, we sum up the derivatives.</p>

<div class="flex-container">
<p>\(\dfrac{\partial cost}{\partial p_1} = \dfrac{\partial cost}{\partial p_2} \cdot \dfrac{\partial p_2}{\partial p_1} + \dfrac{\partial cost}{\partial p_3} \cdot \dfrac{\partial p_3}{\partial p_1} = 2 \cdot (p_4 - y) \cdot w_4 \cdot w_2 + 2 \cdot (p_4 - y) \cdot w_5 \cdot w_3\)</p>
</div>

<p>Don’t worry if it doesn’t sink in immediately. The important thing is to grasp that when this is scaled up, any perceptron’s derivative can be calculated by using the derivatives of the others it feeds into. It’s the forward pass in reverse.</p>

<p>Once we understand this, finding the gradients for all the perceptrons becomes simple, and we can run gradient descent on our MLP.</p>

<div class="flex-container">

<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
</pre></td><td class="code"><pre>    <span class="nx">MLP</span><span class="p">.</span><span class="nx">backward</span> <span class="o">=</span> <span class="nf">function </span><span class="p">(</span><span class="nx">desired</span><span class="p">)</span> <span class="p">{</span>
        <span class="c1">// calculate cost derivatives for the first layer</span>
        <span class="c1">// multiply with activation function derivs</span>
        <span class="kd">var</span> <span class="nx">global_derivs</span> <span class="o">=</span> <span class="nf">cost_deriv</span><span class="p">(</span><span class="k">this</span><span class="p">[</span><span class="k">this</span><span class="p">.</span><span class="nx">length</span><span class="o">-</span><span class="mi">1</span><span class="p">].</span><span class="nx">lastOutput</span><span class="p">,</span> <span class="nx">desired</span><span class="p">)</span> <span class="o">*</span> <span class="k">this</span><span class="p">[</span><span class="k">this</span><span class="p">.</span><span class="nx">length</span><span class="o">-</span><span class="mi">1</span><span class="p">].</span><span class="nf">activation_derivs</span><span class="p">()</span>
        <span class="kd">var</span> <span class="nx">prev_derivs</span><span class="p">;</span>
        <span class="k">for </span><span class="p">(</span><span class="kd">var</span> <span class="nx">i</span> <span class="o">=</span> <span class="k">this</span><span class="p">.</span><span class="nx">length</span><span class="o">-</span><span class="mi">1</span><span class="p">;</span> <span class="nx">i</span> <span class="o">&gt;=</span> <span class="mi">0</span><span class="p">;</span> <span class="nx">i</span><span class="o">--</span><span class="p">)</span> <span class="p">{</span>
            <span class="kd">var</span> <span class="nx">l</span> <span class="o">=</span> <span class="k">this</span><span class="p">[</span><span class="nx">i</span><span class="p">]</span>
            <span class="c1">// clone of this layer to store derivatives</span>
            <span class="kd">var</span> <span class="nx">d</span> <span class="o">=</span> <span class="k">this</span><span class="p">.</span><span class="nx">derivs</span><span class="p">[</span><span class="nx">i</span><span class="p">]</span>
            <span class="k">if </span><span class="p">(</span><span class="nx">i</span> <span class="o">!=</span> <span class="mi">0</span><span class="p">)</span>
                <span class="c1">// store derivs for previous layer if it exists</span>
                <span class="nx">prev_derivs</span> <span class="o">=</span> <span class="k">new</span> <span class="nc">Matrix</span><span class="p">(</span><span class="k">this</span><span class="p">[</span><span class="nx">i</span><span class="o">-</span><span class="mi">1</span><span class="p">].</span><span class="nx">length</span><span class="p">,</span> <span class="mi">0</span><span class="p">)</span>

            <span class="k">for </span><span class="p">(</span><span class="kd">var</span> <span class="nx">j</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="nx">j</span> <span class="o">&amp;</span><span class="nx">lt</span><span class="p">;</span> <span class="nx">d</span><span class="p">.</span><span class="nx">length</span><span class="p">;</span> <span class="nx">j</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
                <span class="kd">var</span> <span class="nx">p</span> <span class="o">=</span> <span class="k">new</span> <span class="nc">Matrix</span><span class="p">(</span><span class="nx">d</span><span class="p">[</span><span class="nx">j</span><span class="p">].</span><span class="nx">length</span><span class="p">,</span> <span class="nx">global_derivs</span><span class="p">[</span><span class="nx">j</span><span class="p">])</span> <span class="o">*</span> <span class="nx">l</span><span class="p">.</span><span class="nx">lastInput</span>

                <span class="nx">d</span><span class="p">[</span><span class="nx">j</span><span class="p">]</span> <span class="o">=</span> <span class="nx">d</span><span class="p">[</span><span class="nx">j</span><span class="p">]</span> <span class="o">+</span> <span class="nx">p</span>
                <span class="nx">d</span><span class="p">[</span><span class="nx">j</span><span class="p">].</span><span class="nx">bias</span> <span class="o">+=</span> <span class="nx">global_derivs</span><span class="p">[</span><span class="nx">j</span><span class="p">]</span>
                <span class="k">if </span><span class="p">(</span><span class="nx">i</span> <span class="o">==</span> <span class="mi">0</span><span class="p">)</span> <span class="k">continue</span>
                <span class="k">for </span><span class="p">(</span><span class="kd">var</span> <span class="nx">k</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="nx">k</span> <span class="o">&amp;</span><span class="nx">lt</span><span class="p">;</span> <span class="nx">prev_derivs</span><span class="p">.</span><span class="nx">length</span><span class="p">;</span> <span class="nx">k</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
                    <span class="nx">prev_derivs</span><span class="p">[</span><span class="nx">k</span><span class="p">]</span> <span class="o">+=</span> <span class="nx">global_derivs</span><span class="p">[</span><span class="nx">j</span><span class="p">]</span> <span class="o">*</span> <span class="nx">l</span><span class="p">[</span><span class="nx">j</span><span class="p">][</span><span class="nx">k</span><span class="p">]</span>
                <span class="p">}</span>
            <span class="p">}</span>
            <span class="nx">global_derivs</span> <span class="o">=</span> <span class="nx">prev_derivs</span>
        <span class="p">}</span>
    <span class="p">}</span>
    <span class="nx">MLP</span><span class="p">.</span><span class="nx">applyDerivs</span> <span class="o">=</span> <span class="nf">function </span><span class="p">()</span> <span class="p">{</span>
        <span class="k">for </span><span class="p">(</span><span class="kd">var</span> <span class="nx">i</span> <span class="o">=</span> <span class="mi">0</span> <span class="p">;</span> <span class="nx">i</span> <span class="o">&amp;</span><span class="nx">lt</span><span class="p">;</span> <span class="k">this</span><span class="p">.</span><span class="nx">derivs</span><span class="p">.</span><span class="nx">length</span><span class="p">;</span> <span class="nx">i</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
            <span class="k">for </span><span class="p">(</span><span class="kd">var</span> <span class="nx">j</span> <span class="o">=</span> <span class="mi">0</span> <span class="p">;</span> <span class="nx">j</span> <span class="o">&amp;</span><span class="nx">lt</span><span class="p">;</span> <span class="k">this</span><span class="p">.</span><span class="nx">derivs</span><span class="p">[</span><span class="nx">i</span><span class="p">].</span><span class="nx">length</span><span class="p">;</span> <span class="nx">j</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
                <span class="k">for </span><span class="p">(</span><span class="kd">var</span> <span class="nx">k</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="nx">k</span> <span class="o">&amp;</span><span class="nx">lt</span><span class="p">;</span> <span class="k">this</span><span class="p">.</span><span class="nx">derivs</span><span class="p">[</span><span class="nx">i</span><span class="p">][</span><span class="nx">j</span><span class="p">].</span><span class="nx">length</span><span class="p">;</span> <span class="nx">k</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
                    <span class="k">this</span><span class="p">[</span><span class="nx">i</span><span class="p">][</span><span class="nx">j</span><span class="p">][</span><span class="nx">k</span><span class="p">]</span> <span class="o">-=</span> <span class="nx">WEIGHT_LRATE</span> <span class="o">*</span> <span class="k">this</span><span class="p">.</span><span class="nx">derivs</span><span class="p">[</span><span class="nx">i</span><span class="p">][</span><span class="nx">j</span><span class="p">][</span><span class="nx">k</span><span class="p">]</span>
                <span class="p">}</span>

                <span class="k">this</span><span class="p">[</span><span class="nx">i</span><span class="p">][</span><span class="nx">j</span><span class="p">].</span><span class="nx">bias</span> <span class="o">-=</span> <span class="nx">BIAS_LRATE</span> <span class="o">*</span> <span class="k">this</span><span class="p">.</span><span class="nx">derivs</span><span class="p">[</span><span class="nx">i</span><span class="p">][</span><span class="nx">j</span><span class="p">].</span><span class="nx">bias</span>
            <span class="p">}</span>
        <span class="p">}</span>

        <span class="k">this</span><span class="p">.</span><span class="nf">clearDerivs</span><span class="p">()</span>
    <span class="p">}</span>
    <span class="kd">var</span> <span class="nx">schema</span> <span class="o">=</span> <span class="p">[</span><span class="mi">1</span><span class="p">,</span><span class="mi">1</span><span class="p">,</span><span class="mi">2</span><span class="p">,</span><span class="mi">3</span><span class="p">,</span><span class="mi">4</span><span class="p">]</span>
    <span class="kd">var</span> <span class="nx">mlp</span> <span class="o">=</span> <span class="k">new</span> <span class="nc">MLP</span><span class="p">(</span><span class="nx">schema</span><span class="p">)</span>
</pre></td></tr></tbody></table></code></pre></figure>

<div>
    <div id="canvas-container-9"></div>
    <button class="reset-button" onclick="currentContainer=9;containers[9].mlp = new MLP([1,[3,'relu'],[3,'relu'],[3,'relu'],4,1])"> Reset </button>
</div>
</div>

<p>And that’s it. We have made a fully operational MLP. We can even change it to a logistic regression network by just changing the last layer’s activation function and the cost, like below.</p>

<p><strong>Note:</strong> You might want to stop it after running because it tends to eat up a lot of computing power and slow down the webpage.</p>

<div class="flex-container">

<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
</pre></td><td class="code"><pre>    <span class="nx">cost_deriv</span> <span class="o">=</span> <span class="nx">neglog_cost_deriv</span>
    <span class="kd">var</span> <span class="nx">schema</span> <span class="o">=</span> <span class="p">[</span><span class="mi">2</span><span class="p">,</span><span class="mi">4</span><span class="p">,</span><span class="mi">6</span><span class="p">,[</span><span class="mi">8</span><span class="p">,</span><span class="dl">'</span><span class="s1">relu</span><span class="dl">'</span><span class="p">],</span><span class="mi">8</span><span class="p">,[</span><span class="mi">1</span><span class="p">,</span><span class="dl">'</span><span class="s1">sigmoid</span><span class="dl">'</span><span class="p">]]</span>
    <span class="kd">var</span> <span class="nx">mlp</span> <span class="o">=</span> <span class="k">new</span> <span class="nc">MLP</span><span class="p">(</span><span class="nx">schema</span><span class="p">)</span>
</pre></td></tr></tbody></table></code></pre></figure>

<div>
    <div id="canvas-container-10"></div>
    <button class="reset-button" onclick="currentContainer=10;containers[10].mlp = new MLP([2,[30,'leakyrelu'],[30,'leakyrelu'],[30,'leakyrelu'],30,[1,'sigmoid']]);containers[10].play=true"> Reset </button>
    <button class="reset-button" onclick="containers[10].play = false"> Stop </button>
</div>
</div>

<p>We will finish this article by exploring some slightly more in-depth problems and solutions that come with an MLP. Since this article is hardly meant to be comprehensive, I recommend consulting more sources after it. There are great lectures by Andrew Ng, Andrej Karpathy and Geoffrey Hinton on Youtube, as well as a good practical series by sentdex.</p>

<h4 id="learning-rate">Learning Rate</h4>

<p>When we were implementing the gradient descent and perceptron algorithms, I glossed over some very important numbers in the code, the learning rates.</p>

<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
4
5
6
</pre></td><td class="code"><pre>    <span class="c1">// 0.01 is for speed purposes</span>
    <span class="nx">xs</span><span class="p">[</span><span class="nx">i</span><span class="p">]</span> <span class="o">-=</span> <span class="mf">0.01</span> <span class="o">*</span> <span class="nx">derivative</span>
    <span class="p">.</span> <span class="p">.</span> <span class="p">.</span> 
    <span class="nx">perceptron</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span> <span class="o">+=</span> <span class="mf">0.3</span><span class="o">*</span><span class="nx">adj</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span>
    <span class="nx">perceptron</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span> <span class="o">+=</span> <span class="mf">0.3</span><span class="o">*</span><span class="nx">adj</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span>
    <span class="nx">perceptron</span><span class="p">[</span><span class="mi">2</span><span class="p">]</span> <span class="o">+=</span> <span class="mf">0.3</span><span class="o">*</span><span class="nx">adj</span><span class="p">[</span><span class="mi">2</span><span class="p">]</span>
</pre></td></tr></tbody></table></code></pre></figure>

<p>
It is important not to just subtract the gradient itself from the weight, as it can lead to jumps that are "too big" and miss the optimum solution. In the simulation below, you can see how a learning rate of 1 works on the graph of \(x^2\), compared to a slightly better 0.9, a reasonable 0.3 and an unreasonable 0.05. Click on the graph to trigger one pass.
</p>

<div class="flex-container">
<div>
    <div id="canvas-container-11"></div>
    <button class="reset-button" onclick="containers[11].xs=[-1]"> Reset </button>
</div>

<div>
    <button class="option-button" onclick="containers[11].xs=[-1];containers[11].a = 1" type="button"> 1 </button>
    <button class="option-button" onclick="containers[11].xs=[-1];containers[11].a = 0.9" type="button"> 0.9 </button>
    <button class="option-button" onclick="containers[11].xs=[-1];containers[11].a = 0.3" type="button"> 0.3 </button>
    <button class="option-button" onclick="containers[11].xs=[-1];containers[11].a = 0.05" type="button"> 0.05 </button>
</div>
</div>

<p>As you can see, a badly picked learning rate like 1 can lead your algorithm to hover above the optimum and never reach it. A slightly less worse rate like 0.9 can keep missing the optimum value by overshooting and take a long time to converge. On the other hand, an unnecessarily small rate like 0.05 wastes time and energy by undershooting. In the end, only experience and educated guesses can get you a good or ideal learning rate, in this case 0.3.</p>

<p>It is useful to know that there are more sophisticated algorithms to manage learning rates, like RMSProp, AdaMax and Adam, which are used almost ubiquitously in practical applications, but are outside the scope of this article.</p>

<h4 id="exploding--vanishing-gradients">Exploding &amp; Vanishing Gradients</h4>

<p>
In backpropagation, we have seen that a perceptron's derivative is calculated by multiplying the derivative of the first layer with weights from the others. In deep networks(networks with many layers), this can present a problem. Let's imagine our weights are all smaller than one, maybe an average of 0.5. In that case, for a perceptron \(x\) layers deep, we get the derivative multiplies by \(0.5^x\), which rapidly goes to zero. We call this a vanishing gradient. It is very harmful for an MLP because a zero derivative means zero gradient, which means zero gradient descent. Vanishing gradients make an MLP freeze and stop improving. 
</p>

<p>
On the other hand, if the weights have an average value of 2, we get our derivative multiplied by \(2^x\), which goes to large values really fast. A ten layer network would have gradient of 1024. You can see why we call them exploding gradients. Just like having a big learning rate makes the MLP skip over optimal weights, exploding gradients make a network unstable and even lead "infinite" weights, which crash the network entirely.
</p>

<p>A solution to both problems is to ensure that the weights are initialized so that the gradients get multiplied by a reasonable number. There are schemes such as normal distributions and Xavier initialization, which you can read about in <a href="https://towardsdatascience.com/weight-initialization-in-neural-networks-a-journey-from-the-basics-to-kaiming-954fb9b47c79">this article</a>.</p>

<h4 id="regularization">Regularization</h4>

<p>Another problem that plagues neural networks is their tendency to overfit. Overfitting is when the model becomes a little <em>too</em> good at predicting the dataset, at the expense of its predictive powers outside it. Think of it like this. When you want to predict a trend, you want a line that goes smoothly in that direction. If you have a lot of micro jumps and fluctuations in the prediction, it calls into judgment how effective it is. Someone saying Apple stock is going up tomorrow after seeing the general trend is more credible than someone who predicts the ups and downs it will have every minute of tomorrow by memorizing every minute it has been on the market so far.</p>

<p>A method commonly used for regularization is ensuring the weights are as small as possible, to avoid huge spikes or unstable behavior. We can do this easily by adding the magnitude of the weight to the cost. <a href="https://towardsdatascience.com/how-to-improve-a-neural-network-with-regularization-8a18ecda9fe3">This article</a> is very good at explaining it.</p>

<h3>Batches</h3>

<p>Using batches in a neural network means adding the derivatives of multiple pieces of data before updating the network’s weights. It gives a marked performance advantage because we can do one backward pass for many forward passes. It also makes sense intuitively because it seems counter-productive to change the entire network based on one piece of data. It sounds like the entire government changing because one person said they didn’t like it. But batches are more like democracy. Every voices their opinions and the most prevalent one is adopted.</p>

<h3>Conclusion</h3>

<p>I hope this article has been helpful in understanding and implementing your own neural networks from scratch. It is far from comprehensive, so I recommend you consult other sources. After all, the search for knowledge is never over.</p>

<p>To infinity and beyond.</p>

<script src="https://cdn.jsdelivr.net/npm/p5@1.1.9/lib/p5.js"></script>

<script src="https://cdn.mathjax.org/mathjax/latest/MathJax.js?config=TeX-AMS-MML_HTMLorMML" type="text/javascript"></script>

<script src="/assets/blog_scripts/mlp_tutorial.js"></script>

<script>
    var i = 0
    var e
    var containers = []

    while((e = document.getElementById("canvas-container-"+i)) && window['container'+i]) {
        var canvas = new p5(window['container'+i], e)
        containers.push(canvas)
        i++
    }
</script>]]></content><author><name>Noel Alemayehu</name></author><category term="programming" /><category term="neural-networks" /><category term="buzz-lightyear" /><summary type="html"><![CDATA[Buzz Lightyear teaches math]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="http://noelnegash.github.io/assets/blog_images/post_covers/buzz-cover.png" /><media:content medium="image" url="http://noelnegash.github.io/assets/blog_images/post_covers/buzz-cover.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Neural Networks From Scratch With Buzz Lightyear (Part 1: Calculus)</title><link href="http://noelnegash.github.io/programming/neural-networks/buzz-lightyear/2020/09/05/buzz-1.html" rel="alternate" type="text/html" title="Neural Networks From Scratch With Buzz Lightyear (Part 1: Calculus)" /><published>2020-09-05T07:40:29+00:00</published><updated>2020-09-05T07:40:29+00:00</updated><id>http://noelnegash.github.io/programming/neural-networks/buzz-lightyear/2020/09/05/buzz-1</id><content type="html" xml:base="http://noelnegash.github.io/programming/neural-networks/buzz-lightyear/2020/09/05/buzz-1.html"><![CDATA[<style> .option-button, .reset-button, .change-button {
    border: none;
    color: white;
    text-align: center;
    text-decoration: none;
    display: inline-block;
    font-size: 16px;
    margin: 4px 2px;
    cursor: pointer;
    position: relative;
    font-family: "Montserrat", sans-serif;
    font-size: 14px;
    line-height: 24px;
    font-weight: 700;
    padding: 12px 35px;
    color: #fff;
    text-transform: uppercase;
    border-radius: 3px;
}
.reset-button {
    background-color: #4CAF50; 
    padding: 15px 32px;
}
.change-button {
    background-color: #4274d6;
    display: block;
    padding: 7px 15px;
    }
.option-button {
    background-color: #a12895;
    display: block;
    padding: 7px 15px;
}
.flex-container {
    display: flex;
    flex-direction: row;
    align-items: center;    
    justify-content: center;
}
.flex-container > * {
    margin: 30px;
}
pre {
   text-align: left !important;
}

h4, h3 {
    font-weight: bold;
    margin: 10px auto;
} </style>

<blockquote>
  <p><strong>Warning</strong>: This is not a comprehensive article on neural networks. While it does discuss the mathematical concepts behind them and expects at least some programming experience, there are many aspects of neural networks that are glossed over. The idea is to create one from scratch with only the bare essential knowledge, leaving the option open for further study on your own. It is meant for people who have tried reading the theory and want to see it simplified or people who want a quick intro which they can build upon later. More theory intensive resources I recommend are the Youtube series by <a href="https://www.youtube.com/watch?v=CS4cs9xVecg&amp;list=PLkDaE6sCZn6Ec-XTbcX1uRg2_u4xOEky0">Andrew Ng</a> and <a href="https://www.youtube.com/watch?v=Wo5dMEP_BbI">sentdex</a>.</p>
</blockquote>

<p>Neural networks are a type of machine learning algorithm that are vaguely inspired by the human brain. They take a certain input and are trained to predict the desired outcome, which could be anything from house prices to pictures of dogs. They are constructed of very simple analogues to neurons, called perceptrons, which are basically linear functions that we will optimize. It’s okay if you don’t understand what all that means. That’s what this article is here for. All that you need to know is that a neural network takes in some numbers and tries to guess the best ones to spit out.</p>

<p>This article will be split into three parts:</p>

<ul>
  <li><strong>Part 1: Calculus</strong>
  An introduction to the prerequisite differential calculus that drives neural networks (with a bunch of interactive animations)</li>
  <li><strong>Part 2: Perceptron</strong>
  Using calculus to code a perceptron, the smallest unit of a neural network (with a bunch of interactive animations)</li>
  <li><strong>Part 3: Multi-Layer Perceptron (MLP)</strong>
  Stringing perceptrons end to end as well as stacking them on top of each other to make more powerful models (with a bunch of interactive animations, obviously)</li>
</ul>

<p>So, without further ado, let’s get into it. Calculus, a terrifying branch of mathematics that’s up there with rocket science and neurosurgery when it comes to difficulty, or so they say. It’s actually very simple. A very good way to see differential calculus is as the slopes of curves. And I will show it to you in the best way to explain slopes, Buzz Lightyear.</p>

<p>We all know his catchphrase “To Infinity and Beyond”. While it’s very cool to say, mathematics tells us that he can never actually reach infinity, let alone go beyond it. Math is cruel like that sometimes. But we can (and do) say that Buzz is approaching infinity the more he goes forward and approaching negative infinity the more he goes backward. And we can measure how fast he is approaching infinity with how steeply he is ascending. In math, we call that the slope/gradient of the line.</p>

<p>In the simple equation ( y = x ), the gradient is 1. And we can set his inclination that way to get a natural looking flight path.</p>

<div class="flex-container">

<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
4
5
6
</pre></td><td class="code"><pre><span class="k">if </span><span class="p">(</span><span class="nb">Math</span><span class="p">.</span><span class="nf">abs</span><span class="p">(</span><span class="nx">x</span><span class="p">)</span> <span class="o">&gt;=</span> <span class="mf">2.5</span><span class="p">)</span>
    <span class="nx">speed</span> <span class="o">*=</span> <span class="o">-</span><span class="mi">1</span>
        
<span class="nx">x</span> <span class="o">+=</span> <span class="nx">speed</span>

<span class="nf">drawBuzz</span><span class="p">(</span><span class="nx">x</span><span class="p">,</span> <span class="nx">x</span><span class="p">,</span> <span class="mi">1</span><span class="p">)</span> <span class="c1">// x, y, and inclination/slope</span>
</pre></td></tr></tbody></table></code></pre></figure>

<div id="canvas-container-0"></div>
</div>

<div class="flex-container">
<div>
<p>
Well, it sounds very obvious when we're talking about a line, but let's imagine other functions, like a parabola \( y = x^2 \). Then it gets harder to talk about his gradient, because it keeps changing. Sometimes Buzz is going up and other times he's going down. That is where derivatives come in. 
</p>
<p>
A derivative of a function is a formula for its gradient based on \( x \). The derivative of any power function \( y = x^n + c \) is given by \( y = n\cdot x^{n-1} \). Note that any constant number added to the expression is ignored. Actually, the derivative of \( y = mx \) itself is given by this formula. And the derivative of \( y=x^2 \) is \( 2x \). We can see this working when we make Buzz's inclination equal \( 2x \) in the animation. 
</p>
</div>
<div id="canvas-container-1"></div>
</div>

<div class="flex-container">
<div>
Another thing to keep in mind is that if the function has a coefficient, we have to keep it as well. So the derivative of \( y = m\cdot x^n + c \) is given by \( y = m\cdot n\cdot x^{n-1} \)
</div>
<div id="canvas-container-2"></div>
</div>

<div class="flex-container">
<div>
And derivatives aren't limited to just power functions. Fore xample, the derivative of \( \sin x \) is \( \cos x \).
</div>
<div id="canvas-container-3"></div>
</div>

<div class="flex-container">
<div>
You can get the derivatives of any combination of functions as well. But you need to follow a certain rule called the chain rule. What it says is that when you are looking for the derivative of a function within a function, you need to calculate the derivatives in layers and finally multiply them.
For example, with the function \( y = \sin 3x \), the derivatives would be \( \cos 3x \) for \( \sin 3x \) and \(3\) for \(3x\), leaving us with \(3 \cdot \cos 3x \) as the derivative for the combined function.
</div>
<div id="canvas-container-4"></div>
</div>

<div class="flex-container">
<div>
Thankfully, addition of functions is very easy to differentiate. We just add the derivatives. For \( \sin 3x + x\), we get the derivative \( 3 \cdot \cos 3x + 1\).
</div>
<div id="canvas-container-5"></div>
</div>

<p>Now you know all the calculus you need to code a basic perceptron. How? By combining it with gradient descent. Gradient descent is the core algorithm behind neural networks. It uses the gradient of a function to find the location at which it has the lowest value. Since an increasing function has a positive gradient and a decreasing function has a negative gradient, in either case, going the opposite direction of the gradient takes you to the minimum of the function.</p>

<p>In the example below, there are 4 Buzz Lightyears scattered on random points in the function, and they all try to go to the minimum using gradient descent. You can use the reset button to make them go all over again. When you do that, you’ll notice that some get stuck on the upper “rung” of the function and think that’s the best they will do. In that case, we say they are caught in a “local minimum”, not the absolute one. You may also notice that some of them disappear over the left edge almost immediately because it keeps going down. Those Lightyears are even closer to the best minimum.</p>

<div class="flex-container">

<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
11
</pre></td><td class="code"><pre><span class="k">for </span><span class="p">(</span><span class="kd">var</span> <span class="nx">i</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="nx">i</span> <span class="o">&lt;</span> <span class="nx">xs</span><span class="p">.</span><span class="nx">length</span><span class="p">;</span> <span class="nx">i</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
    <span class="kd">var</span> <span class="nx">derivative</span> <span class="o">=</span> <span class="mi">3</span><span class="o">*</span><span class="nf">cos</span><span class="p">(</span><span class="nx">xs</span><span class="p">[</span><span class="nx">i</span><span class="p">]</span><span class="o">*</span><span class="mi">3</span><span class="p">)</span><span class="o">+</span><span class="mi">1</span>

    <span class="c1">// subtracts the gradient</span>
    <span class="c1">// 0.01 is for speed purposes</span>
    <span class="nx">xs</span><span class="p">[</span><span class="nx">i</span><span class="p">]</span> <span class="o">-=</span> <span class="mf">0.01</span> <span class="o">*</span> <span class="nx">derivative</span>

    <span class="c1">// recalculate derivative</span>
    <span class="c1">// because x changed</span>
    <span class="nx">derivative</span> <span class="o">=</span> <span class="mi">3</span><span class="o">*</span><span class="nb">Math</span><span class="p">.</span><span class="nf">cos</span><span class="p">(</span><span class="nx">xs</span><span class="p">[</span><span class="nx">i</span><span class="p">]</span><span class="o">*</span><span class="mi">3</span><span class="p">)</span><span class="o">+</span><span class="mi">1</span>
<span class="p">}</span>
</pre></td></tr></tbody></table></code></pre></figure>


<div>
<div id="canvas-container-6"></div>
<button class="reset-button" onclick="resetGradient(6)" type="button"> Reset </button>
</div>
</div>

<p>There is also a variant of gradient descent called gradient <em>ascent</em>. It’s fundamentally the same, except it tries to find the <em>maximum</em> of the function by going along the gradient. Here it is with the same function. Notice that there are local <em>maxima</em> here instead of minima.</p>

<div class="flex-container">


<figure class="highlight"><pre><code class="language-javascript" data-lang="javascript"><table class="rouge-table"><tbody><tr><td class="gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
</pre></td><td class="code"><pre><span class="k">for </span><span class="p">(</span><span class="kd">var</span> <span class="nx">i</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="nx">i</span> <span class="o">&lt;</span> <span class="nx">xs</span><span class="p">.</span><span class="nx">length</span><span class="p">;</span> <span class="nx">i</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
    <span class="kd">var</span> <span class="nx">derivative</span> <span class="o">=</span> <span class="mi">3</span><span class="o">*</span><span class="nf">cos</span><span class="p">(</span><span class="nx">xs</span><span class="p">[</span><span class="nx">i</span><span class="p">]</span><span class="o">*</span><span class="mi">3</span><span class="p">)</span><span class="o">+</span><span class="mi">1</span>

    <span class="c1">// only difference is the plus</span>
    <span class="nx">xs</span><span class="p">[</span><span class="nx">i</span><span class="p">]</span> <span class="o">+=</span> <span class="mf">0.01</span> <span class="o">*</span> <span class="nx">derivative</span>

    <span class="c1">// recalculate derivative</span>
    <span class="c1">// because x changed</span>
    <span class="nx">derivative</span> <span class="o">=</span> <span class="mi">3</span><span class="o">*</span><span class="nb">Math</span><span class="p">.</span><span class="nf">cos</span><span class="p">(</span><span class="nx">xs</span><span class="p">[</span><span class="nx">i</span><span class="p">]</span><span class="o">*</span><span class="mi">3</span><span class="p">)</span><span class="o">+</span><span class="mi">1</span>
<span class="p">}</span>
</pre></td></tr></tbody></table></code></pre></figure>



<div>
<div id="canvas-container-7"></div>
<button class="reset-button" onclick="resetGradient(7)" type="button"> Reset </button>
</div>
</div>

<p>And that’s all there is to it. In the <a href="/programming/neural-networks/buzz-lightyear/2020/09/05/buzz-2.html">next part of this series</a>, we’ll see how we can use gradient descent to make a perceptron, the smallest unit of a neural network, and make it find the best-fit line for arbitrary data.</p>

<script src="https://cdn.jsdelivr.net/npm/p5@1.1.9/lib/p5.js"></script>

<script src="https://cdn.mathjax.org/mathjax/latest/MathJax.js?config=TeX-AMS-MML_HTMLorMML" type="text/javascript"></script>

<script src="/assets/blog_scripts/calculus_tutorial.js"></script>

<script> 
    var i = 0
    var e

    var containers = []
    while((e = document.getElementById("canvas-container-"+i)) && window['container'+i]) {
        var canvas = new p5(window['container'+i], e)

        containers.push(canvas)

        i++
    }
</script>]]></content><author><name>Noel Alemayehu</name></author><category term="programming" /><category term="neural-networks" /><category term="buzz-lightyear" /><summary type="html"><![CDATA[Buzz Lightyear teaches math]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="http://noelnegash.github.io/assets/blog_images/post_covers/buzz-cover.png" /><media:content medium="image" url="http://noelnegash.github.io/assets/blog_images/post_covers/buzz-cover.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">CyberTalents Write-up: Ethiopia National Cybersecurity CTF 2020</title><link href="http://noelnegash.github.io/hacking/hackathon/cybersecurity/2020/08/23/cyber-talents.html" rel="alternate" type="text/html" title="CyberTalents Write-up: Ethiopia National Cybersecurity CTF 2020" /><published>2020-08-23T19:19:18+00:00</published><updated>2020-08-23T19:19:18+00:00</updated><id>http://noelnegash.github.io/hacking/hackathon/cybersecurity/2020/08/23/cyber-talents</id><content type="html" xml:base="http://noelnegash.github.io/hacking/hackathon/cybersecurity/2020/08/23/cyber-talents.html"><![CDATA[<p>Yesterday, I participated in the National Cybersecurity CTF hosted in Ethiopia by CyberTalents. Considering it was my first CTF, as well as my having only one other team member and missing the training sessions given by CyberTalents, I was very satisfied with getting 4th place in the competition out of 37 teams.</p>

<p>The format of the CTF competition was 10 challenges ranging from easy to hard with 7 hours given to solve them. They tested our abilities in branches of cybersecurity such as General Information, Digital Forensics, Web Security, Cryptography and Malware Reverse Engineering. In this article, I will be showing how we solved the 8 questions we figured out during the competition, as well as one we did immediately after. Number 10 was way out of my league. Maybe I’ll write a follow up article when I get it. So without further ado, here are the questions.</p>

<p><strong>Cracker</strong> (<em>General Information</em>)</p>

<p>A simple question, asking us the name of a popular Linux tool used as a “packet sniffer, WEP and WPA/WPA2-PSK cracker”.</p>

<p>The answer is obviously <code class="language-plaintext highlighter-rouge">aircrack-ng</code>.</p>

<p><strong>Unprotected</strong> (<em>Digital Forensics</em>)</p>

<p>In this challenge, we were given <code class="language-plaintext highlighter-rouge">Unprotected.pcap</code>, a packet capture file, and told to find the flag inside it. By opening it in Wireshark, we can see it’s a combination of TCP and HTTP packets.</p>

<div class="flex-container"><img src="/assets/blog_images/2020/08/image-3-1024x928.png" alt="" /></div>

<p>A simple filter <code class="language-plaintext highlighter-rouge">data.data contains "flag"</code> leaves just one TCP packet, which contains the flag <code class="language-plaintext highlighter-rouge">flag{cl3ar_t3xt_15_alway35_5asy}</code>.</p>

<div class="flex-container"><img src="/assets/blog_images/2020/08/image-1024x924.png" alt="" /></div>

<p><strong>Encrypted RSA</strong> (<em>Cryptography</em>)</p>

<p>In this challenge, we are given a RSA-encrypted file called <code class="language-plaintext highlighter-rouge">secret</code>, as well as the <code class="language-plaintext highlighter-rouge">p</code>, <code class="language-plaintext highlighter-rouge">q</code> and <code class="language-plaintext highlighter-rouge">e</code> used to encrypt it. All of this enough information to decrypt the file. So after a fast search on <a href="https://crypto.stackexchange.com/questions/19444/rsa-given-q-p-and-e">crypto.stackexchange.com</a>, we made this script.</p>

<figure class="highlight"><pre><code class="language-python" data-lang="python"><span class="k">def</span> <span class="nf">egcd</span><span class="p">(</span><span class="n">a</span><span class="p">,</span> <span class="n">b</span><span class="p">):</span>
    <span class="n">x</span><span class="p">,</span><span class="n">y</span><span class="p">,</span> <span class="n">u</span><span class="p">,</span><span class="n">v</span> <span class="o">=</span> <span class="mi">0</span><span class="p">,</span><span class="mi">1</span><span class="p">,</span> <span class="mi">1</span><span class="p">,</span><span class="mi">0</span>
    <span class="k">while</span> <span class="n">a</span> <span class="o">!=</span> <span class="mi">0</span><span class="p">:</span>
        <span class="n">q</span><span class="p">,</span> <span class="n">r</span> <span class="o">=</span> <span class="n">b</span><span class="o">//</span><span class="n">a</span><span class="p">,</span> <span class="n">b</span><span class="o">%</span><span class="n">a</span>
        <span class="n">m</span><span class="p">,</span> <span class="n">n</span> <span class="o">=</span> <span class="n">x</span><span class="o">-</span><span class="n">u</span><span class="o">*</span><span class="n">q</span><span class="p">,</span> <span class="n">y</span><span class="o">-</span><span class="n">v</span><span class="o">*</span><span class="n">q</span>
        <span class="n">b</span><span class="p">,</span><span class="n">a</span><span class="p">,</span> <span class="n">x</span><span class="p">,</span><span class="n">y</span><span class="p">,</span> <span class="n">u</span><span class="p">,</span><span class="n">v</span> <span class="o">=</span> <span class="n">a</span><span class="p">,</span><span class="n">r</span><span class="p">,</span> <span class="n">u</span><span class="p">,</span><span class="n">v</span><span class="p">,</span> <span class="n">m</span><span class="p">,</span><span class="n">n</span>
        <span class="n">gcd</span> <span class="o">=</span> <span class="n">b</span>
    <span class="k">return</span> <span class="n">gcd</span><span class="p">,</span> <span class="n">x</span><span class="p">,</span> <span class="n">y</span>
 
<span class="k">def</span> <span class="nf">main</span><span class="p">():</span>
 
    <span class="n">p</span> <span class="o">=</span> <span class="mi">11882546252751469607361356421348933496327112595288260315935663917400681403905188808476289112967043136936045873689827577396206505769293138372274271493958287</span>
    <span class="n">q</span> <span class="o">=</span> <span class="mi">10374751834382966611285517450958269115435289482194774831009591093240922739864785750413607023913149510232252798244495377789107452564252835088008933746132847</span>
    <span class="n">e</span> <span class="o">=</span> <span class="mi">65537</span>
 
 
    <span class="n">cipher_text</span> <span class="o">=</span> <span class="nb">int</span><span class="p">.</span><span class="nf">from_bytes</span><span class="p">(</span><span class="nf">open</span><span class="p">(</span><span class="sh">"</span><span class="s">secret</span><span class="sh">"</span><span class="p">,</span><span class="sh">"</span><span class="s">rb</span><span class="sh">"</span><span class="p">).</span><span class="nf">read</span><span class="p">(),</span> <span class="n">byteorder</span><span class="o">=</span><span class="sh">"</span><span class="s">big</span><span class="sh">"</span><span class="p">)</span>
 
 
    <span class="c1"># compute n
</span>    <span class="n">n</span> <span class="o">=</span> <span class="n">p</span> <span class="o">*</span> <span class="n">q</span>
 
    <span class="c1"># Compute phi(n)
</span>    <span class="n">phi</span> <span class="o">=</span> <span class="p">(</span><span class="n">p</span> <span class="o">-</span> <span class="mi">1</span><span class="p">)</span> <span class="o">*</span> <span class="p">(</span><span class="n">q</span> <span class="o">-</span> <span class="mi">1</span><span class="p">)</span>
 
    <span class="c1"># Compute modular inverse of e
</span>    <span class="n">gcd</span><span class="p">,</span> <span class="n">a</span><span class="p">,</span> <span class="n">b</span> <span class="o">=</span> <span class="nf">egcd</span><span class="p">(</span><span class="n">e</span><span class="p">,</span> <span class="n">phi</span><span class="p">)</span>
    <span class="n">d</span> <span class="o">=</span> <span class="n">a</span>
 
    <span class="nf">print</span><span class="p">(</span> <span class="sh">"</span><span class="s">n:  </span><span class="sh">"</span> <span class="o">+</span> <span class="nf">str</span><span class="p">(</span><span class="n">d</span><span class="p">));</span>
 
    <span class="c1"># Decrypt ciphertext
</span>    <span class="n">pt</span> <span class="o">=</span> <span class="nf">pow</span><span class="p">(</span><span class="n">cipher_text</span><span class="p">,</span> <span class="n">d</span><span class="p">,</span> <span class="n">n</span><span class="p">)</span>
 
    <span class="nf">print</span><span class="p">(</span> <span class="sh">"</span><span class="s">plain textt: </span><span class="sh">"</span> <span class="o">+</span> <span class="nf">str</span><span class="p">(</span><span class="nb">int</span><span class="p">.</span><span class="nf">to_bytes</span><span class="p">(</span><span class="n">pt</span><span class="p">,</span> <span class="mi">128</span><span class="p">,</span> <span class="n">byteorder</span><span class="o">=</span><span class="sh">"</span><span class="s">big</span><span class="sh">"</span><span class="p">))</span> <span class="p">)</span>
 
<span class="nf">main</span><span class="p">()</span></code></pre></figure>

<p>Running the script outputs a series of null bytes followed by <code class="language-plaintext highlighter-rouge">Nice Job, flag is FLAG{Gr3at_J0b_F0r_Th3_D3crypti0n}</code>.</p>

<div class="flex-container"><img src="/assets/blog_images/2020/08/image-5-1024x335.png" alt="" /></div>

<p><strong>Gu55y</strong> (<em>Web Security</em>)</p>

<p>In this challenge, we were given a <a href="http://ec2-18-156-199-115.eu-central-1.compute.amazonaws.com/guessy/&quot;">url</a> to exploit. Its functionality is to take your inputs and store them in your cookies as a serialized php list, so that it may display them again when you refresh the page.</p>

<div class="flex-container"><img src="/assets/blog_images/2020/08/image-6.png" alt="" /></div>

<p>Crossing out SQL injection as a vector, we inspected the the page source and found an HTML comment reading <code class="language-plaintext highlighter-rouge">&amp;lt;!-- I love vim~ --&amp;gt;</code>. Seeing that <code class="language-plaintext highlighter-rouge">vim</code> was involved, we tried to search for known <code class="language-plaintext highlighter-rouge">vim</code> file extensions such as <code class="language-plaintext highlighter-rouge">.php~</code> and <code class="language-plaintext highlighter-rouge">.php.un~</code>. We succeeded in downloading <code class="language-plaintext highlighter-rouge">.index.php.swp</code>, which gave us this code.</p>

<figure class="highlight"><pre><code class="language-php" data-lang="php"><span class="c1">#try to read fl4g.php</span>
<span class="kd">Class</span> <span class="nc">l33t</span><span class="p">{</span>
    <span class="k">public</span> <span class="k">function</span> <span class="n">__toString</span><span class="p">()</span>
    <span class="p">{</span>
        <span class="k">return</span> <span class="nb">highlight_file</span><span class="p">(</span><span class="nv">$this</span><span class="o">-&gt;</span><span class="n">source</span><span class="p">,</span><span class="kc">true</span><span class="p">);</span>
    <span class="p">}</span>
<span class="p">}</span></code></pre></figure>

<p>This code excerpt gave us two things, the location of the flag and the vector with which we could get it. We made a <code class="language-plaintext highlighter-rouge">l33t</code> object and set its <code class="language-plaintext highlighter-rouge">source</code> to <code class="language-plaintext highlighter-rouge">fl4g.php</code>. We then serialized it in a list, url-encoded it and put the value in our <code class="language-plaintext highlighter-rouge">list</code> cookie.</p>

<p><strong>serialized</strong>: <code class="language-plaintext highlighter-rouge">a:1:{i:0;O:4:"l33t":1:{s:6:"source";s:8:"fl4g.php";}}</code></p>

<p><strong>urlencoded</strong>: <code class="language-plaintext highlighter-rouge">a%3A1%3A%7Bi%3A0%3BO%3A4%3A%22l33t%22%3A1%3A%7Bs%3A6%3A%22source%22%3Bs%3A8%3A%22fl4g.php%22%3B%7D%7D</code></p>

<div class="flex-container"><img src="/assets/blog_images/2020/08/image-7-1024x488.png" alt="" /></div>

<div class="flex-container"><img src="/assets/blog_images/2020/08/image-8-1024x488.png" alt="" /></div>

<p>Refreshing the page gives us the flag <code class="language-plaintext highlighter-rouge">flag{5w337_PHP_0bj3c7_!nj3c7!0n}</code></p>

<p><strong>Habibamod</strong> (<em>Digital Forensics</em>)</p>

<p>In this challenge, we were given another packet capture file, <code class="language-plaintext highlighter-rouge">Habibamod.pcap</code>. In Wireshark, we can see it’s an HTTP session where a file is being uploaded.</p>

<div class="flex-container"><img src="/assets/blog_images/2020/08/image-9-1024x826.png" alt="" /></div>

<p>The contents of the file are a JSON object with properties <code class="language-plaintext highlighter-rouge">data</code> and <code class="language-plaintext highlighter-rouge">encoder</code>.</p>

<p><code class="language-plaintext highlighter-rouge">{"data": ".!…!!..!!.!!…!!….!.!!..!!!.!!!!.!!.!.!.!…!..!!.!.!….!!.!.!.!…!…!!..!!!!..!..!!…..!!!.!.!.!!!..!..!…!…!!..!.!.!!…!!..!!…..!!..!…!!..!.!..!!…..!!..!!…!….!.!…….!!.!!!..!!..!….!.!!!..!..!..!.!!!..!!.!…….!!.!!.!.!!….!..!!.!!!.!!.!..!.!!.!!!..!!..!!!.!!!!!.!", "encoder": "ZGVmIG15X2VuY29kZXIoZGF0YSk6CiAgICBiaW5fcmVwID0gJycuam9pbihmb3JtYXQob3JkKGkpLCAnYicpIGZvciBpIGluIHgpCiAgICByZXR1cm4gYmluX3JlcC5yZXBsYWNlKCcwJywnLicpLnJlcGxhY2UoJzEnLCchJykgCg=="}</code></p>

<p>Decoding <code class="language-plaintext highlighter-rouge">encoder</code> with base64, we get a python function called <code class="language-plaintext highlighter-rouge">my_encoder</code>. After analyzing it, we made a python function of our own to reverse it. We figured out <code class="language-plaintext highlighter-rouge">encoder</code> was base64 because of the telltale <code class="language-plaintext highlighter-rouge">==</code> at the end.</p>

<figure class="highlight"><pre><code class="language-python" data-lang="python"> <span class="k">def</span> <span class="nf">my_encoder</span><span class="p">(</span><span class="n">data</span><span class="p">):</span>
    <span class="n">bin_rep</span> <span class="o">=</span> <span class="sh">''</span><span class="p">.</span><span class="nf">join</span><span class="p">(</span><span class="nf">format</span><span class="p">(</span><span class="nf">ord</span><span class="p">(</span><span class="n">i</span><span class="p">),</span> <span class="sh">'</span><span class="s">b</span><span class="sh">'</span><span class="p">)</span> <span class="k">for</span> <span class="n">i</span> <span class="ow">in</span> <span class="n">x</span><span class="p">)</span>
    <span class="k">return</span> <span class="n">bin_rep</span><span class="p">.</span><span class="nf">replace</span><span class="p">(</span><span class="sh">'</span><span class="s">0</span><span class="sh">'</span><span class="p">,</span><span class="sh">'</span><span class="s">.</span><span class="sh">'</span><span class="p">).</span><span class="nf">replace</span><span class="p">(</span><span class="sh">'</span><span class="s">1</span><span class="sh">'</span><span class="p">,</span><span class="sh">'</span><span class="s">!</span><span class="sh">'</span><span class="p">)</span> </code></pre></figure>

<figure class="highlight"><pre><code class="language-python" data-lang="python"> <span class="k">def</span> <span class="nf">our_decoder</span><span class="p">(</span><span class="n">data</span><span class="p">):</span>
    <span class="c1"># the previous function converts a string to a bitmap 
</span>    <span class="c1"># and then changes 0 and 1 to . and ! respectively
</span>
    <span class="c1"># reverse the string replacement
</span>    <span class="n">data</span> <span class="o">=</span> <span class="n">data</span><span class="p">.</span><span class="nf">replace</span><span class="p">(</span><span class="sh">'</span><span class="s">.</span><span class="sh">'</span><span class="p">,</span><span class="sh">'</span><span class="s">0</span><span class="sh">'</span><span class="p">).</span><span class="nf">replace</span><span class="p">(</span><span class="sh">'</span><span class="s">!</span><span class="sh">'</span><span class="p">,</span><span class="mi">1</span><span class="p">)</span>
    
    <span class="c1"># change the bitmap to numbers
</span>    <span class="n">data</span> <span class="o">=</span> <span class="p">[</span><span class="nf">int</span><span class="p">(</span><span class="n">data</span><span class="p">[</span><span class="n">i</span><span class="p">:</span><span class="n">i</span><span class="o">+</span><span class="mi">8</span><span class="p">],</span> <span class="mi">2</span><span class="p">)</span> <span class="k">for</span> <span class="n">i</span> <span class="ow">in</span> <span class="nf">range</span><span class="p">(</span><span class="mi">0</span><span class="p">,</span> <span class="nf">len</span><span class="p">(</span><span class="n">data</span><span class="p">),</span> <span class="mi">8</span><span class="p">)]</span>

    <span class="c1"># change the numbers to a string
</span>    <span class="n">data</span> <span class="o">=</span> <span class="sh">''</span><span class="p">.</span><span class="nf">join</span><span class="p">([</span><span class="nf">chr</span><span class="p">(</span><span class="n">i</span><span class="p">)</span> <span class="k">for</span> <span class="n">i</span> <span class="ow">in</span> <span class="n">data</span><span class="p">])</span>
    <span class="k">return</span> <span class="n">data</span></code></pre></figure>

<p>Putting <code class="language-plaintext highlighter-rouge">data</code> into this function gives us the flag <code class="language-plaintext highlighter-rouge">Flag{TMCTFy0urDec0de0f!@nd.Is@ma7ing}</code></p>

<div class="flex-container"><img src="/assets/blog_images/2020/08/image-10.png" alt="" /></div>

<p><strong>GoldASM</strong> (<em>Malware Reverse Engineering</em>)</p>

<p>In this challenge, we are given an assembly file <code class="language-plaintext highlighter-rouge">GoldASM.asm</code>. It described a function that had many repetitions of this format.</p>

<figure class="highlight"><pre><code class="language-assembly" data-lang="assembly">    mov     rax, QWORD PTR [rbp-24]
    mov     eax, DWORD PTR [rax]
    cmp     eax, 70
    jne     .L2
    mov     rax, QWORD PTR [rbp-24]
    add     rax, 4
    mov     eax, DWORD PTR [rax]
    cmp     eax, 76
    jne     .L2</code></pre></figure>

<p>We realized that the code was going across a string and checking its value against hardcoded numbers. By removing the repeated parts of code and making the rest human readable, it becomes obvious.</p>

<figure class="highlight"><pre><code class="language-assembly" data-lang="assembly">    0:
    compare     eax, 70   # 'F' char
    4:
    compare     eax, 76   # 'L' char</code></pre></figure>

<p>There were certain spots though, where it wasn’t simply comparing. For example, one line doubled the number before comparing it and another subtracted from it.</p>

<figure class="highlight"><pre><code class="language-assembly" data-lang="assembly">    8:
    add eax, eax
    compare eax, 130     # meaning we want 65, or 'A'
    24:
    subtract eax, 75
    compare eax, 2       # meaning we want 77, or 'M'`</code></pre></figure>

<p>There were also some points which required a byte be less than a number instead of equal. After dealing with these edge cases, we got a list of numbers <code class="language-plaintext highlighter-rouge">[70, 76, 65, 71, 123, 95, 75, 51, 101, 98, 95, 48, 110, 95, 83, 104, 49, 110, 105, 110, 103, 95, 125]</code>. Converting them into characters gave us the flag <code class="language-plaintext highlighter-rouge">FLAG{_K3eb_0n_Sh1ning_}</code>.</p>

<p>Sadly, I don’t have any screenshots of the other challenges because they were web based and CyberTalents took them down when the hackathon ended.</p>

<p>All in all, this hackathon was lots of fun and and I’m interersted in doing it again if I can. I’ll make sure to document the whole thing too.</p>]]></content><author><name>Noel Alemayehu</name></author><category term="hacking" /><category term="hackathon" /><category term="cybersecurity" /><summary type="html"><![CDATA[Writeup for a cybersecurity hackathon]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="http://noelnegash.github.io/assets/blog_images/post_covers/ctf-cover.png" /><media:content medium="image" url="http://noelnegash.github.io/assets/blog_images/post_covers/ctf-cover.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>