This document says how Automatic Ruby is put together: what the pieces are, what each is responsible for, which way they depend on each other, and how a value travels through a run.
What the system is for belongs to REQUIREMENTS.md. The
rules a change is held to belong to POLICY.md. The two public
contracts — the Recipe format and the plugin interface — are specified in
PLUGINS.md; this document explains the machinery that implements
them and does not restate them.
It stands on its own. Nothing in it is completed by a document kept in another repository.
Automatic Ruby is a general-purpose composition framework, not host infrastructure and not a purpose-built application. The design therefore protects the small set of contracts that make composition possible without treating every framework implementation detail or every shipped plugin as equally permanent.
- The framework is the small part. It loads a Recipe, finds classes by name, and calls them in order. It has no domain knowledge, and gaining some would be a design error rather than a feature.
- Plugins are the large part, and they are replaceable. They are found on a search path, not registered in a list, so that adding one touches no framework file and removing one leaves nothing dangling.
- One value connects everything. Because every plugin takes and returns the same shape, composition needs no adapter, no type negotiation and no schema.
- The core is loadable without its plugins' dependencies.
require 'automatic'pulls in the framework and nothing a plugin needs. Each plugin requires its own libraries, at the top of its own file, so that an absent gem is an error only for a Recipe that asked for that plugin. - The library never exits and never prints. Exit status is decided by the entry point; user-facing text is written by the entry point or logged.
Maintenance strength follows the architectural layer:
- Core contracts are strongly protected. The Recipe format, plugin contract, single pipeline shape, lookup and override semantics, execution order and established CLI behaviour are depended on outside this repository.
- Framework internals may improve inside those contracts. The loader, CLI, helpers and internal structure are not frozen merely because they are old, provided the invariants and public contracts remain intact.
- Plugins are intentionally replaceable. They may be added, repaired, replaced or removed as their external systems change. A plugin whose service or interface no longer exists is not preserved by simulation merely to retain catalogue size.
Composability is an architectural invariant, not an optimization preference. It is the architectural property that defines the framework: a small core, independent plugins, one pipeline representation and Recipe-level composition. A change that replaces those properties changes the identity of the system and is judged as an architecture change.
A typical run is a short, linear flow. Markdown is one publisher at the edge, not a second pipeline representation:
source plugins -> shared pipeline -> optional filters / stores -> publishers
bin/automatic process entry point; exit status only
|
v
lib/automatic/cli.rb option parsing, subcommands, error reporting
|
+---------------------------+
| |
v v
lib/automatic.rb lib/automatic/opml.rb diagnostic helpers
(run, directories) lib/automatic/feed_parser.rb
|
+-------------------+
| |
v v
lib/automatic/recipe.rb lib/automatic/pipeline.rb
| |
| v
| Automatic::Plugin::* plugins/<category>/<name>.rb
| | ~/.automatic/plugins/<category>/<name>.rb
v v
lib/automatic/log.rb lib/automatic/feed_maker.rb
| lib/automatic/http.rb
v
standard output
Dependency points downward, and there is no edge back up:
bin/automaticknows onlyAutomatic::CLI.Automatic::CLIknows the framework. Nothing in the framework knows the CLI.Automatic::Pipelineknows how to find and call a plugin. It knows no plugin.- A plugin knows
Automatic::Log,Automatic::FeedMaker,Automatic::FeedParser,Automatic::Httpand its own libraries. It knows no other plugin. Automatic::LogandAutomatic::FeedMakerare leaves. They depend on nothing in this repository.
Automatic::Recipe and Automatic::Pipeline do not know each other. Both are
driven by Automatic.run.
The process entry point, and deliberately almost empty. It puts the
installation's lib on the load path, requires automatic/cli, and exits with
the status Automatic::CLI.run(ARGV) returns.
It contains no option definition, no subcommand and no policy. Everything a test would want to exercise is therefore in a library file, reachable without spawning a process.
Everything that belongs to being a command:
- Builds the
OptionParser:-c/--config,-h/--help,-v/--version. - Holds the subcommand table (
scaffold,unscaffold,autodiscovery,feedparser,inspect,opmlparser,log) as a hash of name to callable. - Implements
scaffoldandunscaffold, which are the only filesystem operations the framework performs on its own behalf. - Catches the framework's own exceptions, prints one line to standard error, and returns the exit status.
CLI.run returns an Integer and never calls exit. That is what lets a
spec assert on a status and on captured output instead of a subprocess.
Exit status is decided in exactly one place:
| Status | Meaning |
|---|---|
0 |
The Recipe ran, or the subcommand did its work, or help or version was printed |
1 |
The run or the subcommand failed, or no work was requested |
2 |
The command line was rejected by the option parser |
The libraries a subcommand needs are required inside that subcommand, not at the
top of the file, so running a Recipe loads neither feedbag nor the OPML
parser. The framework's own requires still apply: require 'automatic' pulls in
the Recipe loader, the pipeline, the log and the two feed adapters, and through
them Ruby's rss.
The module itself, holding the two directories the rest of the system resolves paths against, and the one method that runs a job:
root_dir— the installation root, set by the entry point.user_dir—~/.automatic, or an override whenAUTOMATIC_RUBY_ENV=test, which is how the specs point the loader at a fixture directory without writing to a real home directory.plugins_dir,config_dir,user_plugins_dir— derived from those two.run(recipe:, root_dir:, user_dir:)— sets the directories and hands the Recipe toPipeline.run.
It requires the framework's own files and nothing else. It does not require a
plugin, and a plugin's dependency never appears here.
Bundler setup for a source checkout: if the repository's Gemfile is present,
set it up so bin/automatic resolves the locked gems.
It is a convenience for running from a checkout, not a requirement. When the gem
is installed normally there is no Gemfile to find and this file does nothing,
which is the correct behaviour: an installed library must not impose a bundle on
the program that requires it.
Turns a Recipe file into an object the pipeline can iterate.
- Resolves the path: a bare name is looked for in
~/.automatic/configfirst, and anything not found there is treated as a path as given. - Parses the YAML safely — the document may contain only the plain types a
Recipe needs, so a Recipe cannot name a Ruby class to instantiate. Aliases are
permitted, because they are a legitimate way to share a block of settings
between plugins. This is a second line of defence, not the trust boundary; see
REQUIREMENTS.mdsection 17. - Wraps the result in
Hashie::Mash, which is why a plugin entry answers toplugin.moduleandplugin.config, and why an absent setting reads asnilrather than raising. - Applies
global.log.leveltoAutomatic::Log. This is the onlyglobalkey the framework reads. - Exposes
each_plugin, which yields the entries ofpluginsin order. - Raises
Automatic::InvalidRecipeErrorfor a document that is not a mapping, or that carries no usablepluginssequence.
It performs no validation of a plugin's own config. Whether a setting is
required, and what it must look like, is the plugin's business, because the
framework cannot know.
The core, and the shortest file that matters.
load_plugin(module_name) resolves a class name to a file:
module_name.underscoreturnsSubscriptionFeedintosubscription_feed.- The category directories are listed,
~/.automatic/plugins/*before<root>/plugins/*, so the user directory wins. - For each, if the directory's own name is a prefix of the underscored module
name, the remainder is the file name: directory
subscriptionand modulesubscription_feedgivesubscription/feed.rb. - The first existing file wins, and is registered with
Automatic::Plugin.autoload, so the file is read when the constant is first used. - Nothing matched raises
Automatic::NoPluginErrornaming the module.
The category directory is therefore not a label: it is half of the lookup key. This is what lets a plugin be added by dropping in a file, and a shipped plugin be overridden by putting a file of the same name in the user directory.
run(recipe) is the pipeline:
pipeline = []
recipe.each_plugin do |plugin|
mod = plugin.module
load_plugin(mod)
klass = Automatic::Plugin.const_get(mod)
pipeline = klass.new(plugin.config, pipeline).run
endEach plugin's return value is the next plugin's input. There is no branching, no
inspection of the value between steps, and no exception handling: a plugin that
raises ends the run, for the reason given in REQUIREMENTS.md section 12.
A Logger on standard output behind a level filter.
Log.level(name)sets the threshold, fromglobal.log.level.Log.puts(level, message)emits whenlevelis at least the threshold.- Levels are
info,warn,error,none, compared by their position in that list;noneis the threshold that admits nothing. - Both a
Stringand aSymbolare accepted for a level, because plugins have always passed both, and an unknown level is treated asinforather than raising in the middle of an unattended run.
It is a module with state rather than an injected object. That is a consequence
of plugins calling Automatic::Log directly, which keeps a plugin's signature
to (config, pipeline).
The adapters between "some data" and the pipeline shape, and the one way in for what is fetched.
FeedParser.get_url(url)fetches a URL and parses it as a feed.FeedParser.parse_html(html)builds a feed whose items are the page's links, which is how the link and Tumblr subscription plugins work.FeedMaker.generate_feed(hash)builds one item-like object from plain values.FeedMaker.create_pipeline(items)builds one feed object from a list of them. Any plugin producing items from a non-feed source ends with this call.FeedMaker.content_provide(url, data)builds a one-item feed carrying an arbitrary payload incontent_encoded, which is the route by which the XML subscription plugin feeds the Fluentd provide plugin.
The first two use Ruby's bundled rss library. That is the reason the pipeline
value has the shape it has.
Http.read(url)fetches a URL and returns the body andHttp.open(url)yields the stream, for a caller — an HTML parser, which detects a page's encoding for itself — that would rather not be handed a decoded string;Http.uri(url)returns a validated URI andHttp.fetchable?(url)answers whether there is one.
Automatic::Http exists because the decisions a fetch implies — which schemes
are allowed, how long to wait, how many redirects to follow, what to send as a
User-Agent — were being made separately by every plugin that fetched, mostly by
omission. It is a helper of about twenty lines and not a client: a plugin that
wants something else calls Ruby directly. The scheme allowlist is the part that
earns it a file of its own, because a link in a pipeline item comes from a feed
and URI.open on such a string will read a local file as readily as an
article.
Every plugin is a class in Automatic::Plugin, constructed with (config, pipeline) and answering run. The contract is specified in
PLUGINS.md section 3; what matters to this document is how the
categories divide responsibility:
| Category | Directory | Receives | Returns | Role |
|---|---|---|---|---|
Subscription |
subscription/ |
usually an empty pipeline | a pipeline | Acquire from outside |
CustomFeed |
custom_feed/ |
usually an empty pipeline | a pipeline | Build a feed from a non-feed source |
Filter |
filter/ |
a pipeline | a pipeline | Select, reorder or rewrite |
Store |
store/ |
a pipeline | a pipeline, usually reduced | Persist, and drop what was seen before |
Provide |
provide/ |
a pipeline | the same pipeline | Emit the payload elsewhere |
Notify |
notify/ |
a pipeline | the same pipeline | Send a notification |
Publish |
publish/ |
a pipeline | the same pipeline | Send the result out, or print it |
The categories are a convention with one mechanical consequence — the directory
name is part of the lookup key (section 4.6) — and no other. Nothing enforces
that a Filter does not reach the network.
Replaceability is part of this boundary. A plugin is not given the same permanence as the Recipe format or the plugin contract itself. The framework preserves the rules that let plugins compose; it does not preserve a plugin whose external purpose has disappeared, and it does not move a plugin's domain behaviour into the core merely to make that behaviour permanent.
Publish is the boundary at which the pipeline meets a representation that is
not the pipeline's. A publishing plugin reads the value described in section
4.8, writes it out in whatever form its destination wants — a line on a
terminal, a request body, a record for a log collector, a document on disk — and
returns the value it was given, unchanged. The conversion is the plugin's whole
job and it happens in one direction, at one point:
RSS-shaped pipeline
|
v
Publish<Something> the plugin: serialize, then emit
|
v
the destination's own form a terminal line, a JSON record, a Markdown document
PublishMarkdown is one of these and is architecturally unremarkable: it
renders each item as a Markdown section and writes the result to standard output
or to a file. What matters to this document is where it is not:
- Not in the framework.
Automatic::Pipelinegains no knowledge of Markdown,Automatic::FeedMakergains no Markdown constructor, and no framework file mentions the format. A reference to it inlib/would be the same design error as a reference to any other plugin (section 3). - Not a second pipeline value. Markdown is produced from the pipeline at
the moment the pipeline ends; it is never passed along one. The plugin returns
its input, so a plugin placed after it receives exactly what it would have
received without it, and Invariant 2 of
POLICY.mdis untouched. - Not implicit. Nothing appends it to a Recipe.
Pipeline.runruns the entries the Recipe lists, in order, and that is still the whole of its behaviour.
Its specification — what it writes for each field, how HTML in a body is
reduced, where the output goes — is in PLUGINS.md section 6.7,
because it is a plugin's specification and not a property of the design.
One shared piece sits inside plugins/ rather than in lib/, because it is
plugin implementation and the framework does not use it:
plugins/store/database.rb— theAutomatic::Plugin::Databasemixin: opens the SQLite database named in the Recipe, creates the table from the including class'scolumn_definitionwhen it is absent, and providesfor_each_new_feed, which yields only items whose key is not already stored.StorePermalinkandStoreFullTextare this mixin plus a model and a column list.StoreDigesttakes the database part of it and decides for itself what has been seen, because it is identified by a digest of an item's content rather than by the linkfor_each_new_feedreads.
Fallbacks inside the installation, used when the corresponding part of the user
directory is absent: db/ for SQLite files, config/ for the example Recipes
that scaffold copies out, assets/ for data files a plugin needs.
automatic -c feed2console.yml
|
| CLI parses the command line, resolves the recipe path
v
Recipe.new(path)
| reads YAML safely, wraps in Hashie::Mash
| sets the log level from global.log.level
v
Automatic.run(recipe:, root_dir:)
| sets root_dir and user_dir
v
Pipeline.run(recipe)
|
| pipeline = []
|
| entry 1: module SubscriptionFeed
| load_plugin -> plugins/subscription/feed.rb
| SubscriptionFeed.new(config, []).run
| FeedParser.get_url(each configured feed)
| -> [feed]
|
| entry 2: module FilterIgnore
| load_plugin -> plugins/filter/ignore.rb
| FilterIgnore.new(config, [feed]).run
| drops items matching a keyword
| FeedMaker.create_pipeline(kept items)
| -> [feed']
|
| entry 3: module StorePermalink
| load_plugin -> plugins/store/permalink.rb
| StorePermalink.new(config, [feed']).run
| opens ~/.automatic/db/<db>, creates the table if absent
| for_each_new_feed: skips links already stored, inserts the rest
| -> [feed''] containing only what had not been seen
|
| entry 4: module PublishConsole
| PublishConsole.new(config, [feed'']).run
| prints each item
| -> [feed'']
v
the final value is discarded; CLI returns 0
Three properties of that flow are the design:
- The pipeline narrows. Subscription plugins produce, filters and stores reduce, publishers consume. A Recipe is normally read in that order.
- A store plugin is what makes a Recipe safe to run every five minutes. It is the only thing standing between the operator and a duplicate.
- Nothing between the steps inspects the value. The framework never looks inside a feed.
The last entry decides what the run leaves behind, and changing it changes
nothing else. The same four steps ending in PublishMarkdown produce a document
instead of terminal output:
acquire SubscriptionFeed, SubscriptionLink, CustomFeed*, ...
| many sources, one shape
v
filter FilterIgnore, FilterSanitize, FilterSort, ...
|
v
store StorePermalink: what has not been seen before
|
v
publish PublishMarkdown: serialize at the boundary
|
v
Markdown a file, or standard output
|
+--> read by a person
+--> searched with grep, processed with the usual tools
+--> committed, diffed, kept
+--> handed to a program, a language model or an agent
Everything above the last box is unchanged from the previous diagram, which is the point: the acquisition, the filtering and the de-duplication are the same plugins over the same value, and the exit is the part the Recipe chooses.
Two kinds, and they do not mix.
Framework settings come from the Recipe's global mapping, and there is one
that is read: global.log.level. global.timezone and global.cache appear in
the shipped examples and in Recipes in the wild, and nothing reads them. They
are kept — removing them would break nothing but would edit operators' files for
no gain — and they are documented as inert in
PLUGINS.md section 2.4 so that no one adds a behaviour to them
by accident.
Plugin settings are the config mapping of a plugin entry, passed to the
constructor and read by nobody else. The framework does not validate, default,
merge or type-check them. A plugin that needs a value it was not given decides
what to do, and the common choice — treating an absent retry or interval as
zero — is why nil.to_i appears throughout the plugins.
There is no environment-variable configuration, with one exception:
AUTOMATIC_RUBY_ENV=test permits the user directory to be overridden, and
exists for the tests.
The layers handle errors differently, and on purpose.
Plugins own transient failure. A plugin that reaches the network wraps the
attempt, logs the failure, and retries retry times with interval seconds
between attempts. What a plugin must not do is swallow an error and return a
value that looks like success.
The framework owns nothing it cannot fix, so it catches nothing. Its own failures are named:
| Exception | Raised when |
|---|---|
Automatic::NoRecipeError |
Pipeline.run was given no Recipe |
Automatic::NoPluginError |
No file was found for a module named in a Recipe |
Automatic::InvalidRecipeError |
The Recipe is not a mapping, or has no plugins |
All three derive from Automatic::Error, so a caller can rescue the framework's
failures without rescuing everything.
The CLI is the only place that turns an exception into a message and a
status. It reports the framework's errors as one line on standard error and
returns 1. Anything else it lets propagate, because an unexpected exception is
a defect and its backtrace is wanted.
- A library file never calls
puts. It logs. Automatic::CLIwrites to standard error for diagnostics and standard output for requested output — help, version, and the results of the diagnostic subcommands.- Publishing plugins whose entire purpose is to print (
PublishConsole,PublishConsoleLink,PublishMarkdownwith nofilesetting) write to an output object held in an instance variable, defaulting to$stdout. That is what lets their specs assert on what was printed by substituting a double.
The log and a publishing plugin's output therefore share standard output, which
has never mattered — a pretty_inspect dump and a log line interleave without
either becoming unreadable — and matters for a plugin whose output is a document
meant to be redirected into a file. The design does not resolve this by giving
the log a second destination: Automatic::Log writing to standard output is
long-standing behaviour that Recipes and cron entries are written around
(REQUIREMENTS.md section 14), and changing it for one plugin's benefit would
be the framework acquiring a plugin's problem. It is resolved where the choice
already exists — in the Recipe:
global.log.level: nonesilences the log for that run, leaving standard output carrying the document alone;- or the plugin writes to a file it is given, and the log keeps standard output.
Both are stated in PLUGINS.md section 6.7 and in
DEPLOYMENT.md, which is where an operator looks.
The design choices that exist for the tests:
CLI.runreturns a status instead of exiting, so command-line behaviour is a unit test.Automatic.user_dir=accepts an override underAUTOMATIC_RUBY_ENV=test, so the plugin loader can be pointed atspec/user_dirand the user-directory precedence rule can be asserted.spec/user_dir/plugins/store/mock.rbexists for exactly that.- The plugin contract is a constructor and one method with no ambient input, so
a plugin spec is: build a pipeline, construct,
run, assert on the result.AutomaticSpec.generate_pipelinebuilds the pipeline. - Each plugin requires its own libraries in its own file, so a plugin whose gem is not installed is one spec that does not load, rather than a suite that does not start.
The default suite covers the framework and the plugins that need neither the
network nor a credential. What that leaves out, and why, is in
PLUGINS.md section 6 and in the README's testing section.