Showing posts with label DEVOPS. Show all posts
Showing posts with label DEVOPS. Show all posts

Saturday, September 19, 2020

Dependency Management Best Practices

Every single technology project uses dependencies. No matter the language you use, you always will use dependencies. Monoliths and Monorepos make dependency management more manageable by having it all in the same place. Services and Microservices make them a bit more challenging because now they are spread across each server, updating libs is a challenge that requires automation to be fixed. Besides automation, some best practices are also needed. Several times engineering teams take dependency management for granted. After all, we are just adding some strings and numbers to a file, right? So how complicated can that get? Often dependency management issues only appear after time and scale. Dependency Management is a super important and relevant discipline. Today I want to share some best practices to make your life better and make sure you scale your codebase with speed and solid practices rather than piles of tech debt and pain. So In case your not geek enough, the post icon image is a famous Death Star for Star Wars, which also becomes popular by the Death Star diagram from microservice organizations like Amazon, Netflix, Twitter, and other companies. 

Dependency Management Best Practices

Often Best practices are depending on contexts. Once you have complexity several times, best practices do not necessarily translate from one company to another; however, there are cases where they apply and make sense, not always, but for Dependency management IMHO, these are Golden Rules and excellent practices to be followed. Here are some Dependency Management best practices: 

  • 1. Use a Dependency management tool (ant, maven, gradle). Do Explicit Dependency management.
  • 2. Use Artifact Management Solution (Nexus, Archiva, Artifactory) - Management but Mainly: Cache.
  • 3. Remove dependencies you dont use.
  • 4. Use Consistency Versioning (MAJOR, MINOR, SEC/PATCH).
  • 5. Do not use multi-project EXTERNAL POMS.
  • 6. Keep Dependencies up to Date (But does not update at deploy time - Immutable Infrastructure)
  • 7. Use Dependencies Carefully (Shared-Libs) avoid coupling as much as you can.

Dependency management is not only about using some specific or better tools. It's about a process and culture which requires attention and automation as well.  Let's do a deep dive into each of these practices and understand why they are important. 

Use the Dependency management tool (ant, maven, gradle). Do Explicit Dependency management.

It might sound obvious but it's not uncommon to see infrastructure projects downloading binaries manually and not using explicit dependency management via tools like Ant/ivy, Maven, and Gradle. No Matter the language you use, no matter if is an engineering or DevOps code, you should do explicit dependency management. Because is easier to maintain and we can relly on common process and tools for improvements and housekeeping. 

Explicit dependency management means, explicit defining dependencies on a file that is used by the dependency management tool. You should avoid having embedded dependency and scripts who download dependencies outside of your main tool. 

Use Artifact Management Solution (Nexus, Archiva, Artifactory) - Management but Mainly: Cache

Software tends to grow with your business. As you grew build times can get very slow. A cache is a must-have feature. Because there are multiple engineers downloading artifacts from the web and you often have multiple cloud environments like DEV, STAGING, STREE, PROD, etc... Dependency management solutions like Nexus, Archiva, Artifactory, also can help with better dependency management but one of the main benefits is to have a central cache and central repository management/location. 

Remove dependencies you dont use.

This might sound silly. But un-used dependencies are bad as Dead Code because they make upgrade efforts harder and they end up creating technical debt. It's not uncommon that dependencies have 3rd party dependencies too and you might be dependency from a 3rd party dep instead from a direct dep and have that scenario for a dep you dont use is bad. It's unclear, confusing, raise false positives, and makes reasoning about refactoring efforts much much harder. 

Use Consistency Versioning (MAJOR, MINOR, SEC/PATCH).

Versioning is something old as the snakes in the jungle(like we use to say in Brazil). However people still dont get it right. Why? People know when to use MAJOR(Major API Breaking change), Minor(Minor change no breaking backward compatibility), and Security/Patch release (minor bug fixe or security patch, not impact). But people do not do it. Why? Most of the time is a combination of lack of discipline and lack of ownership and pain of upgrading people dependencies. Central teams can be great in sense of reducing some costs but certainly, they hide some of the pains that if people would face them directly the would definitely deal with the problem differently.  Having consistent versioning is super important, for Design, for Testing, and for health engineering practice I would argue. This part requires discipline and every single binary should embrace this principle. 

Do not use multi-project EXTERNAL POMS.

Don't be fooled by the word POM. This principle works for any dependency management tool. You should not share multi-project external configs for dependency management. Either you have a monolith or monorepo where you have all code in one place or if you do have multiple github repositories you should not share these files(poms). Because? Well because they are EVIL. They create coupling they make upgrades harder and they kill microservices. 

If you will have shared libraries they should be:

 * Small

 * Independent 

 * Isolated (dont have poms, not share configs)

Otherwise, you will build a distributed monolith and binary coupling will prevent you from upgrade when you need it. Never trade coupling for convenience or developer experience.  However, if you have a monolith or a monorepo is perfectly fine to share poms. 

Keep Dependencies up to Date (But does not update at deploy time - Immutable Infrastructure)

Another super important practice is to keep your dependencies updates. Thats important for several reasons such as:

  * Prevent Bugs

  * Fix Security Bugs

  * Reduce Tech Debt

Update Dependencies often works with the same principles as Branches in Configuration Management. If you gonna have a long-lived branch(which you should avoid at all costs) you need to do merges every day so it reduces the complexity and issues on an old fashion bing bang boom merge. Libraries updates work in the same way. For minor, security patches, even minors should be able to upgrade it easily. 

DevOps has a principle called - Immutable Infrastructure, you do not want to upgrade libs before doing a deploys or when a service restart. Because that breaks the principle of immutable infrastructure. However, at the same time, you want to AUTOMATE your dependency management and update libs frequently. Often engineers do not have the mindset to keep updating libs, which can be fixed with proper plugins and automation. 

Use Dependencies Carefully (Shared-Libs) avoid coupling as much as you can.

When we ship internal shared libraries we need to be very careful. Shared Libs should be treated by applying the same principles we apply for Services. It's super important to pay extra attention to 3rd party deps in shared libs in order to avoid binary coupling. It's fine to use shared-libs, sometimes thats the right solution, however, there is a huge abuse of internal shared libs on the technology industry. 

Better Dependency management helps Design and Testing. It makes CI/CD more effective and in the long-run increases the speed and ability to ship better and more frequent software. Currently, we live in an era where every company is trying to do proper CI/CD, Observability, Services, DevOps, SRE, and many other important matters however we often forget dependency management is an important sub-part of Building that end ups charging a high price at scale. 

Cheers,

Diego Pacheco

Thursday, September 17, 2020

Github cli tool (gh)

gh is the new  Github command-line tool. I'm super excited about this new tool. It makes the engineers / devops engineer experience much better. It allows you to use github(not fully yet but the basics) in the command line, so you dont need to use the browser. GH is written in GO and I believe in the future we will have amazing tooling from Github to CI/CD using GitHub Actions. So let's get started.


The Video

The Instructions

Cheers,

Diego Pacheco

Tuesday, September 1, 2020

Linux Terminal Goods V

It's time for a new kick-ass console application to boost your productivity. The tools can be useful both for engineering productivity on your local workstation/laptop or doing some remote pair programming but also for the cloud as you profiling or debugging something. So if you did not check it out the other 4 posts please check it out here: I, II, III, IV. So like Bruce Buffer would say it Iiiiiiiiiiiiiiiiiiiiiiiiiiiiiiiiiiit's time! Linux Terminal Goods V!


Diskonaut

Kubernetes and Docker are great, however, time to time the fill your 500GB HDD.So In order to save space, please don't delete your Steam games. Use diskonaut and figure out the big files and delete them pretty sweet and easy.

Httpstat

Latency can be a Bitch sometimes. Httpstat helps you to figure it out. :-) 

Delta

This is not an air company. However delta helo you to figure out Diffs pretty on the fly. There is even a github like a template so you even will forget that you are in a terminal. 

Webify

Sometimes you want to test things pretty dirty and easy, other times you want just to have fun, mostly if you just want to have fun, check it out webify. It all any BASH command or script be turned into a web server. So let's pretend it's secure and just have a good time with Webify.  

Ranger

ls can be deprecated now. There is something much cooler and with a badass name that's Ranger. Ranger allows you to navigate thought the file system via terminal and open vim at the end. 

That's it for today! 

Cheers,

Diego Pacheco

Monday, August 24, 2020

Hardening Production

Often successful companies have significant growth. Meaning: grow in structure. More people, more departments, more managers, more coordination and often more overlaps. Enterprise companies always end up having some form or governance issue. I remember at the beginning of the SOA world "Governance" was a very bad word and there was abuse. After 3 waves of Agile manifestation, we are with much less "Governance" look like we still have plenty of abuse. Abuse can be also called WASTE(Lean Concept). Some rules can really promote best practices and better software like: Do not share your internal data stores, Expose data via Common Interfaces like (HTTP / gRPC).  Systems tend to grow and get more complex as companies grow with them. Every single improvement(refactoring) often means more investment($$$) which also can be saying as business entropy as the time passes it only gets worst(more complex and more expensive). When we analyze refactorings individually(business impact) they might not make sense but that cannot be the only measure otherwise entropy will kill you in the long run, so there is the complex chess math going on. Bad code or bad design is not the only source of entropy, tests can be big offenders too. 

Unit Tests have can have Waste

People often value unit tests as they are the best kind of testing technique. IMHO that happens for a couple of reasons. One because there was a lack of unit tests in legacy applications, often legacy web apps and desktop apps. So it's natural to say unit test all the things. For the frontend, there are better ways to test and have value like double down on linters and static tying like TypeScript for instance. Also for the frontend, you can double down on Snapshot Testing and tools like Jest. However, backend Unit Tests have more value and are more important. We need to understand that unit tests can have waste, meaning: Flaky tests that break all the time for bad practices, coupling with external things like database IDs, or just because of mocking abuse. One thing companies always want to do is raise the bar, which makes sense however often that bar means more coverage. 

All metrics can be gamed 

That's the issue, as you as for more coverage the goal became to produce more tests, not often means produce better tests or make sure we dont have waste in tests. Whatever number you get 100%, 90%, 80%, 70% you might have WASTE but higher numbers have a higher chance to have waste there. Should we have no metrics for them? well, thats a hard problem. Since this is subjective is really hard to measure and really hard to have a clear safe a universal signal. 

The Issue with Mocks

The issue is people end up abusing from mock. Either you have the real thing and you are doing a real integration or E2E test or you want to test something else and the mocks are a way to isolate your code from other stuff so you can test in isolation. That is fine, however, the issue is most of the time people are testing mocks. There is no value in testing mocks. Mock is a form of coupling if you have a shared library mock that could be a huge step into migrations and patches. 

More Diversity (CI/CD, Observability, More types of tests) 

IMHO there are other things we can do than just (add more coverage) like have a continuous CI/CD pipeline working which will increase ownership and by having the code being potentially deployed at any time(Readiness) you increase the health of the codebase by saying you need to keep the build and test passing all times. The issue with Release Trains or Release calendars is often there are components you are not deploying and people end up no caring and creating a state of build failures and tests that do not pass. With CI/CD there is no such thing. No having CI/CD means you will afford much more waste which is bad. 

Observability and the ability to understand what's going on is also a huge improvement and observability is also a form of testing. So it's another dimension that can be explored. Finally, I would say you can look for having more test diversity like Mutation Testing, Property-Based Testing, Snapshot Testing, Chaos Engineering & Stress Testing. These other forms of testing allow you to do more with less and also uncover issues you might not catch with regular unit/integration tests. 

Hardening Production the whole point 

Finally, I would say it does not matter if works in your machine. The whole point is to work in production. Working in production means, testing in production, because able to deploy software there and test with real traffic and real production hardware without impacting the user experience, so we need to have isolation and automation mechanisms in prod in order to deal with the then. We need to understand that production is a shared responsibility and everyone should be holding accountable in regards to production. We want to harden production because is where the value is and where we deliver value to customers. The best place to be improved is the production. 

Cheers,
Diego Pacheco

Monday, August 3, 2020

Terraform Module

Terraform is a great DevOps tool. Terraform has support for Custom Module which is a killer function. Functions allow you to take much more advantage of Terraform and avoid the right custom code in bash or other languages. Today I made a video showing some functions in terraform like how we can work Custom Module. So Let's get started.


Video


Code


Cheers,
Diego Pacheco


Terraform Null-Resource & Local-Exec

Terraform is a great DevOps tool. Terraform has support for Dynamic Provisioning which is a killer function. Functions allow you to take much more advantage of Terraform and avoid the right custom code in bash or other languages. Today I made a video showing some functions in terraform like how we can work Dynamic Provisioning. So Let's get started.


Video


Code


Cheers,
Diego Pacheco

Terraform Foreach

Terraform is a great DevOps tool. Terraform has support for Dynamic Resources which is a killer function. Functions allow you to take much more advantage of Terraform and avoid the right custom code in bash or other languages. Today I made a video showing some functions in terraform like how we can work Dynamic Resources. So Let's get started.


Video


Code


Cheers,
Diego Pacheco

Terraform Functions

Terraform is a great DevOps tool. Terraform has support for functions which is a killer function. Functions allow you to take much more advantage of Terraform and avoid the right custom code in bash or other languages.  Today I made a video showing some functions in terraform like how we can work with external JSON files and use functions. So Let's get started.


Video


Code


Cheers,
Diego Pacheco

Monday, June 8, 2020

Linux Terminal Goods IV

Every software, DevOps, architect, QA engineer loves terminal applications. Console applications make uses more productive than GUI ones. One big important aspect of console/terminal application is the fact that they are scriptable and can be orchestrated with other automation solutions. This post is the sequence of the saga Linux Terminal Goods, If you did not check out previous posts you can do it here: Episodes I, II and III. I don't know if you realize but most of new and cool Linux console apps are written in Rust now a days.  So Let's get started.





fd


fd is a simple, fast, user-friendly alternative to find.



sd


sd is a intuitive, find and replace, similar to sed but much better.



tokei

tokei allow you to count your lines of code(Loc) pretty easily.



ytop

ytop is a like htop but much better.



git-fuzzy

git-fuzzy is really cool. Basically they enhanced git experience using fzf.

git fuzzy status


git fuzzy log


I hope you like it.

Cheers,
Diego Pacheco

Monday, May 11, 2020

DevSecOps: Are we reducing silos now?

DevOps, as movement and set of principles, did a great job making Operations and development the same integrated thing pretty much. However, industry-wide implementations are not quite there. There is an overlap with SRE(Site Reliability Engineering), and quickly you find DevOps Engineerings, DevOps Architects, DevOps Directors, DevOps Managers. There are plenty of DevOps departments out there. The same noise happened with Agile, where the company structures do not change, and silos still exist. Now we are about to extend the reach of DevOps to Security. Belive me of not the naming does not bother me much. There is a DevSecOps manifesto.

DevSecOps Manifesto

Leaning in over Always Saying "No"
Data & Security Science over Fear, Uncertainty, and Doubt
Open Contribution & Collaboration over Security-Only Requirements
Consumable Security Services with APIs over Mandated Security Controls & Paperwork
Business Driven Security Scores over Rubber Stamp Security
Red & Blue Team Exploit Testing over Relying on Scans & Theoretical Vulnerabilities
24x7 Proactive Security Monitoring over Reacting after being Informed of an Incident
Shared Threat Intelligence over Keeping Info to Ourselves
Compliance Operations over Clipboards & Checklists


DevSecOps make sense; There are lots of exciting and sensible principles there. We need to work closely with security and improve things; however, I want to point out other aspects and challenges here. 

The issues we need to consider are:
  1. Principles do not matter effect.
  2. What happens with the pre-requisites? 
  3. What happens for low-regulated industries and companies? 
  4. Is DevSecOps really about security? 
  5. Are we reducing silos and making organizational changes now?

So Let's get started and go aspect by aspect, down to the rabbit hole. Sorry to disappoint you if you thought this would be a security-related post.

1. Principles do not matter effect

Pop Quiz! What Agile, Lean, SOA, Microservices, DevOps have in common? 

A) All Buzzwords people dont understand but say they do it.
B) You say you want to do it, but you do not know what it means.
C) Last +10 years of hype-cycles? 
D) All the above

Albert Einstein once said, "I don't need to know how to answer the sum of two numbers, but how the sum principle works." That's the issue. We have this action about tools and processes, and we forget the principles. We applied some principles as before as still call our selves agile, lean, DevOps, Microservice, DevOps, and now we have DevSecOps to add.

2. What happens with the pre-requisites? 

Another question. Can we implement DecSecOps ops without implementation DevOps properly? Can we implement DevOps properly without implementing Agile properly? Can we implement DevOps properly without Microservices? Can we deploy Microservices properly before knowing SOA? One movement is evolution and input into the next move, sure. But is all about evolution? Are we missing something big here? 

Do we have any pre-requisite to doing DevSecOps? I always see DevSecOps material related to CI/CD and Microservices. Can we reduce the attacker blast radius with a monolith application and globally shared database across all the monoliths? 

3. What's happens for low-regulated industries and companies? 

So IMHO, DevSecOps is not only about high-regulated industries that need to deal with PII data like ensure-tech, fin-tech, banks, and such. Like Vogels said once recently, "Everybody needs to be a security expert now." IF you have a digital product you need to worry about security, Data Breaches are more popular than ever, and deftly could kill your brand and user experience. Unfortunately, what happens is that security is often the last thing that people care about it. There was a time that Tests was the last thing to be done in a project, now is security. 

4. Is DevSecOps really about security? 

Yes, but not limited. It's also about proper architecture, cloud engineering, DevOps, Agile, and several other things. Might you say I dont need to worry about anything I can focus just on security? So how do we?

  1. Deploy security patches to all machines, services without downtimes?
  2. How we Update configurations at runtime without downtime at the massive number of microservices? 
  3. How do we reduce the attacker and reliability Blast radius? Are they opposite after all? 
  4. How do we rotate Kubernetes Cluster Keys? Can we turn that keys? Are we using HTTPS? How do we know all this?
  5. How do we validate best practices at scale? Do we need to change the Build process, or we do all manually? 

Quickly all these questions should require infrastructure, automation, and architecture around lots of non-security related topics like CI/CD, Automation, Observability, Better Deploy Pipelines, Dynamic Configuration Systems, Discovery, Automated Remediations and much more.

5. Are we reducing silos and making organizational changes now?

So this is the big question right. Are we going to continue to do the same security anti-patterns like store passwords in github? Store credentials in the EC2 machines instead of a Secrets Manager like Vault? Do JDBC SQL code without Proper Binding and allow SQL-Injection? After all and declare victory and say we are doing "DevSecOps." The whole industry needs to LEARN and need to change. Both things are hard, but at some point, things need to change. 

Cheers,
Diego Pacheco

Sunday, May 10, 2020

SRE 101

SRE is pretty how right now. Several companies are interested in but do people really know what it means and how it can be done? Lots of people asked me about what I think about SRE and what's the relation with DevOps. So I record this video and hopefully will help you to understand what it means and grasp some of the basic concepts and practices.  If you are doing DevOps correctly you are doing SRE. SRE is like the DBA in the 90s easily one of the most important folks on the company keeping the business running and make sure users have a good experience in sense of no service disruption or no slowness. So let's get started!




Video



Sides




Cheers,
Diego Pacheco

Saturday, April 4, 2020

Observability & Domain Observability: From Understanding to Value

Observability is a must-have a property any mature distributed system solution and/or digital product.  There is no way to "buy" observability, you need to earn it. Observability is like car insurance. No one likes to pay insurance, however, if there is a care crash you do really really want it. However, you cannot acquire the insurance at the very moment of the crash it does not work that one. The car insurance metaphor would work partially If you have a monolith system. As you have microservices, therefore distributed systems, you have many more failure points and in order to detect that the "car crashed" could be much more hard and complicated. There are many challenges in order to have observability on the system. The main idea is that Observability is something LIVE which means it is not a one time job. So you want to have a better understanding of your system via instrumentation which will lead to better observability and this loop keep going. Some time about I was blogging about Telemetry and microservice, so you might want to check this out here 1, 2 and 3.

 Observability Pillars

Observability is about Metrics, Logs, Traces, Dashboards, Alerts, and Testing. Logs are important for troubleshooting or debugging, however, you do not want to start with logs.  You won't start with alerts and dashboards; In order to have good alerts you need to know a couple things:

* What to monitor?
* What is an anomaly?
* How to fix it?

If you create alerts without these 3 simple questions is very likely you will have noise alerts, false positives and therefore alerts will not help you out. Dashboards are nice because you can spot patterns and during troubleshooting, you can correlate things. As I said before Observability should not be done in a waterfall approach, so you can start with something and improve as you go.

Traces are a great tool, traces are much easier to collect in comparison with custom metrics, which will require application/service instrumentalization.  Traces are good to see where things are getting slow considering you have several microservices downstream calls.

Domain Observability

Domain Observability is quite new. The idea is to instrument the application with domain/business metrics not only system or technology metrics. However, domain observability is a new tern the idea of sending business metrics to centralized telemetry solution is quite old.

Why domain observability is important? Because how do we know the business is being effective? How do we know that the customers are using the product? Opening some UIs more than others? So Domain observability has lots of relation with A/B Testing, Split Traffic, and other pipeline patterns.

Observability is about understanding the system and whats is going on, Domain observability is about understanding the business. Any serious product discovery person doing real customer science need to have domain observability.

SRE and Proactive Work

SRE is about keeping the water flowing into pipes. As plumbing is sometimes a few people care or see it if your site is down everybody will see and complain. Not only you will disrupt your customer's experience, loose money and even sometimes harass your brand. Digital products need to care about SRE and reliability.

Due to the Digital Transformation initiatives and waterfall approaches, deploys in production are often delayed and therefore Observability is delayed.  At some point it makes sense, why would you pay for something in production since you are not deploying anything there right? Right, however, observability is not only about "incidents" or "reliability".

Have you ever feel your development is getting slow, having issues to see where the errors are? Well, you might be lacking some Stability Mindset and practices.  You might have tests and still take a lot of time to figure out where the issues are this will slow down your team productivity.  Having said so, Observability is not a DevOps, Architecture, Product Discovery only but also an engineering practice in order to increase reliability of the system.

From Understanding to Value

Observability is about understanding, understand what is going on, understand how your system behaves, accelerate your troubleshooting when you are coding new features, understand how your customers are interesting with your features not only to prioritize features but also to learn more about your customers to deliver better solutions.

In that sense Observability shifts from a very technical thing to a business concern, central do deliver VALUE in Digital products. In order to deliver value to your customers, you need to understand, deliver as fast as possible, improve the solutions as you go. Observability and Domain Observability need to prioritize and taken seriously.  Have you ever start accounting how much money you lose with incidents? How much money you lose by hunting bugs? How much money you lose by searching where the issue is in your code?

Instrument the system, Observe, Understand and Repeat! That's the way to go.

Cheers,
Diego Pacheco

Tuesday, March 31, 2020

Pipelines From CI/CD to Self Operating Systems

Several years ago everybody was talking about and trying todo deploy automation. This is a fixed problem, right? well, several companies now have good pipelines but most of the tech industry is not quite there. There are several reasons why release pipelines are more important than they look. First of all, Deploy != Release, I talked about that on my previous post about Testing in Production. Several companies have doubted if they need the newest forms of CD(Continuous delivery) like GitOPS but I can say almost 100% sure, this is a must-have. So let's step back and talk about Lean and Kanban and business agility. Even if your business is not asking to increase your deploys and release frequency you should be in pursuit of that. Lean has the concept of LEAD TIME, DevOps is about Lean if you have doubts go read the amazing Phoenix project from Gene Kim. Lead time is about from Idea to Production and how long it takes to get there. Production is the best place to be, so you must increase both frequencies of deploy/release and this means going to production as much as possible.

Why Production Matters?

Production is the best place to be! There are some features that can only be understood in production via A/B/n Testing and with real Domain Observability. You might say, that's too much for me or I dont have basic testing in place why should I increase my deploys frequency? Experiments are super important for the Discovery track, some features can only be understood with real user traffic and several rounds of experiments. Experiments might go way beyond background color, CSS style and play with words to see if we sell more, dont get me wrong thats value on that. However, some experiments require backend changes and therefore you need to have different versions of you backend services somehow, either via backward/forward compatibility or via multiple versions of your services running in prod.

Besides business experiments, there is something called growth. Growths teams become super popular on SV startups and big sites like Facebook. Growth teams are the modern sales and marketing departments. Guess what are the important tools they use? Experiments, A/B/n Testing and other user experience and discovery tools. Right, that's not enough argument for you? ok.

There are companies who have several legal and compliance changes that they need to release pretty quickly and if that does not happen it means losing money and even having to be sued by customers. Right, that's still not enough reason? They think this way, imagine we are playing a Brazillian football game imagine 2 hypothetical scenarios where A) You just score 1 goal per game. B) You won't score more goals you just make sure the ball past mid-field. Not increasing deploy frequency is playing defense and playing defense is super dangerous. Not increasing deploy frequency is like piling up balls on the midfield but zero goals, what's the sense of that?

From CI to CD

Continuous integration started as a good idea but thanks to Git Flow it got messed up. Git Flow Sucks. There are several reasons, first of all, it's because is too complex and error phone, second because it is a smell that shows you architecture is fragile. Git flow becomes popular because of mobile development with often had pretty bad architecture and lack of modularity and isolation so there was several people had to work on same files.

CI was about to have the code being "integrated" all the time. This means all good in the trunk/master branch, so no long live branches, long-lived branches are the opposite of continuous integration. Wow but I have a Jenkins server doing builds all the time, it does not matter because the code is all over different branches.


Here we have a very basic and overall CI pipeline. First of all the engineer code on his local machine and do several commits on git and at some point, he pushes the changes to git. Them a Jenkins pipeline is a trigger, it could be scheduled to run every 5 minutes of after every push, IMHO after every push is much better. Once the code is checkout on Jenkins machine the build phase will be kicked in, so here depending on your language several things will happen like, linters will be run in case of frontend development with angular / react, in case of backend development like Java, Go, Rust, Scala, the code will be compiled and unit tests, snapshot tests, mocked contract tests will run. Some advance pipelines also run Integration Tests, E2E tests and sometimes stress tests. After tests pass reports will be generated in the sense of coverage and test results and the binaries will be deployed often to an application server, HTTP server or s3.

CD is the next level, which stands for continuous Deployment. It's absolutely important to understand Continuous deployment can be done with almost the same tools like CI, like Jenkins, Database automation tools like liquibase or flyway, unit tests. However, Cd requires more things like Feature Toggles / Canary and Split traffic.

Feature Toggles / Canary and Split Traffic

Here is where things get interesting since you will be deploying a lot to the production you need to be able to enable and disable features and releases and also to automatic release or rollback changes in order to scale this process without getting crazy. Split traffic needed in order to reduce risk and make sure we dont affect the User experience. 

Canary means automation and one of the superiors and final levels of pipeline maturity. It's important to have proper observability in place(Centralized Logs, Dashboards, Alerts and better instrumentation on the applications and microservices to send data to the canary solution). A good canary implementation will compare several metrics dimensions like Cpu Usage, memory, network, HTTP errors, disk usage, latency percentiles like p75, p90, p99, 99.5 and much more. Netflix does that with more than +500 metrics.  So once you have all this data you can allow some threshold of decay or say you need to be improving by 1% or whatever it makes sense of that service, and therefore this should be configured per services bases.


CD has a more complete pipeline, It's required to have automation for your infrastructure, like databases, cases, full-text search engines, messaging, load balancers and any other component your services relly on. Often infrastructure deployment happens because of service deployment and happens in separate pipelines. Often this work is done using tools like Terraform, Ansible, Cloud Formation, etc...

After the infrastructure deploys them you will push your code and run build and test like you would do in your CI pipeline, the biggest difference for me is not the fact that Stress Test / Load Tests are mandatory but what will happen is, first of all, DEPLOY IN PRODUCTION. I'm sure some folks shat their pants just be reading deploy in production. Being a grown-up means doing hard things. Deploy in production should be a normal thing, not a myth, not an "special" access that only a DevOps team can be doing (If there is a devops team there are no devops and most likely you talent tensity is low).

Once you are in production you want a test in production, reply real user traffic and as all the candy comparisons go well you want rollout to your users, blue-green is the basics, whats you want to do is release to 1% them 5%, 10%, 25%, 50%, 75, 90%, 100% of your users so if you get something wrong you won't disrupt the majority of your users.

Having Better Pipelines

It's easy to mess up with pipelines, here are some tips and lessons learned that will help you out to build better pipelines:

* Do not use Git Flow, try simple models like CI/CD or even GitHUB Flow.
* Avoid multiple downstream jobs, because they increase pipeline complexity.
* Keep an eye on Test performance, make sure you dont have flaky and slow tests, make it all parallel.
* Avoid do obscure and steps that your team does not understand. Stick to the basics and slowly improve.
* Do not hard code credentials, use Hashicorp Vault.
* Make sure you have notifications in place, email and slack are important.
* Avoid long scheduling cycles like 1 time per day.
* Consider dedicated pipelines for the infrastructure tasks like create, search, upgrade, destroy clusters.
* Try to use Terraform instead of Cloud Formation because it is much better and simple.
* Use Tests for Terraform specs.
* Make sure you tune up Jenkins agents and use proper instances size, avoid t2 medium.
* Have more than 1 agent or you will have bottlenecks
* Consider using spot instances and shutdown Jenkins overnight and weekend to save costs.
* Automate your Jenkins service too :-)
* Make engineers responsible for their pipelines and not a "DevOps" team responsibility
* Create templates and guides to help your engineers improve productivity.

GitOPS

GitOPS makes lots of sense if you are running workloads in Kubernetes, it does not matter if is on Ec2 with Kops or with a managed service like EKS or GKE. The easiest way to go EKS on AWS is with EKSCTL by the way.  So GitOPS means that all kubernetes cluster spec changes go to GIT, via Pull request which is super great, it does not matter if you use plain yaml specs or if you are using helm charts(luckily helm v3 because <= v2 its a bit ugly).


GitOPS is not that different from CI/CD, there are just a few differences, first all it's all trigger by Git hooks via Pull Requests. In your build phase besides the application build, you also need to build the docker image or any other container runtime or packing format you might use, in the future maybe we will all be using WASM.  Them you run your tests normally and when you will deploy, first of you need to deploy the cluster changes and laters on the application changes. Using kubernetes canary, Feature Toggles, Split Traffic is much easier to do it so.

Self Operating, Self Healing, Self Remediation Systems

The next level which goes beyond GitOPS is self-operating, self-healing, self-remediation systems. Thinking like a Tesla car who auto-drives, auto-parks for you. You might want to check out some of my experiences with these systems in my previous blog posts:

* Experiences Building a Cassandra Orchestrator(CM)
* Lessons Learned using AWS Lambda as Remediation System

The whole idea is simple. If you think trought its not crazy at all. AWS  and GCP offer more and more services like that, called so "Managed Services" like RDS, ElastiCache, S3, MSK, EKS, and the list go on and on.

If self-operating, self-healing,self-remediation systems are the future why not everybody does it? First of all it all depends on your use cases and scale. If you can you should use managed services if you use another open-source solution and is super important to your business and you have lots of use cases on it them it would make sense to build a self-operation, self-healing, self-remediation system or internally managed services for short. For microservices, GitOPS should be enough for you, for the DATA layer you might need to go one level down.

Improve Business Agility

Teams are agile, agile has +20 years, now we need to go to the next step, Organizations need to be agile they need to have agile goals, Think more in Impact and Outcome instead of output/feature for product discovery and there is no sense to do that without improving the reliability and time to market. Therefore having better pipelines is a must-have investment you need to do and will pay off in the long term.

Cheers,
Diego Pacheco

Chuyên mục văn hoá giải trí của VnExpress

.

© 2017 www.blogthuthuatwin10.com

Tầng 5, Tòa nhà FPT Cầu Giấy, phố Duy Tân, Phường Dịch Vọng Hậu, Quận Cầu Giấy, Hà Nội
Email: nguyenanhtuan2401@gmail.com
Điện thoại: 0908 562 750 ext 4548; Liên hệ quảng cáo: 4567.