Showing posts with label statistics and diagrams. Show all posts
Showing posts with label statistics and diagrams. Show all posts

Monday, October 07, 2013

Picture this: Venn diagram comparing 4 world religions

I have been trying to improve my Venn diagramming skills so that I could get to the point of creating this graphic comparing the New Testament, the Qur'an, the Tao Te Ching, and the Analects of Confucius. I can still imagine a good number of improvements, but it now gets across the basic picture. (Click the image to view a larger version.)
This shows the results of the word clouds / word prevalence studies done previously for the New Testament, the Qur'an, the Tao Te Ching, and the Analects of Confucius. There could be room for improvement in using something more sophisticated than the freeware word cloud tools that were available at the time, in the translations chosen, the methods, and the diagramming skills. Still, first attempts have to start somewhere.

The methods: I started with all the words in each document that were used at more than a certain frequency, percentage-wise. When I began to diagram those I realized it might give something of a false picture to include a word for one document but not another, if a word fell just barely below the inclusion threshhold in the second document. So the final "prominent words" diagram is based on anything above the prominence threshhold in any of the documents, but allows a word to be recognized as important in other documents if it made the list of top 100 words in other documents. I realize that it might be better to expand the diagram to include more words, but frankly my diagramming skills aren't yet ready for the next level of complexity. 

The diagram: The most commonly-used word in each document is bolded: God, Allah, Tao, and Master. This decision to bold only one word is an area for improvement. In some documents the most-used word decisively dominates over the other common words, where in other documents this is not the case. For instance, in the New Testament "Jesus" is mentioned very nearly as often as "God". If a word is very prominent in one document, but at the lower end of common words in another document, currently the graph does not show that. So to take this diagram to the next level, the graph needs to give an idea of the relative prevalence of a word in each document.

Interpreting the results: Because of the translation and culture differences,some caution is called for when interpreting the results. For instance, we wouldn't want to assume that the Tao has no interest in "Truth" just because it's in the other bubbles but is not in the Tao's. What if a translation difference would have given "truth" more prominence, or what if "knowledge" has the same place in the Tao that "truth" has in some of the other documents? Likewise with some of the other differences; there are often equivalents.

Choosing the documents: The New Testament and the Qur'an represent the two most widely-believed faiths in the world. Some would say the Old Testament might be included for Christianity; the Hadiths (at least some collections) might be included for Islam. But since I have to start somewhere, I start with these documents and simply make sure that everyone is aware of what is and is not yet included. The other documents -- the Analects and the Tao Te Ching -- are prominent and respected in their own right, but some people might rather have seen other documents chosen for comparison. When making the decisions which to include for this very first diagram, one deciding factor was having a reasonably clear choice of "first texts" to analyze. This made the Analects of Confucius more of a logical starting point than, say, Buddhism, where the question "Which documents to study as the first foundation?" is a far trickier question. To be clear, time permitting in this life, I would hope to expand the study to include more faiths. But my diagramming skills are currently challenged enough for 4, so this will do for today. If anyone is impatient for my progress, I'd invite them to help and diagram other works themselves.

Some real similarities: All four share "call(ed), heaven(s), people, word(s)" -- something that could indicate a shared emphasis on a message to or from or about heaven, for people. However, none of these are the most prominent word in any of the documents. There is less similarity here than we have seen in some earlier diagrams.

Some recognizable groupings: The theistic religions focus on God (see the overlap area between the New Testament and the Qur'an); it's also the only time when two of these documents share the same most common word, if we allow "God" and "Allah" as the same. The Tao and the Analects share a focus on government. The New Testament and the Analects show more of an awareness of personal relationships and interactions: ask and reply, disciples, father and son, and love. The Qur'an and the Tao share common words "fear" and "follow", where "fear" is a prominent word in the Qur'an, and in the top 100 words in the Tao; and where "follow" is prominent in the Tao and in the top 100 words in the Qur'an. (Here is a time where seeing the relative prevalence of the words would show that the Qur'an and the Tao have fairly little in common.)

Some areas of uniqueness: The Tao Te Ching focuses on the Way (Tao) as its single most important concept. The Analects focuses on the master, Confucius, propriety and superiority -- along with some other noteworthy persons besides Confucius. The New Testament has a unique emphasis on Jesus Christ and the spirit, along with Paul, and "brothers" or the concept of brotherhood. The Qur'an has a focus on the concept of a "book" -- something of a stand-in for the idea of revealed religion. It also gives prominence to "penalty" (punishment) and "fire" (hell) in a way that is not found in any of the others. None of these observations will be surprising to anyone who has read the works under consideration.

The point: The differences and similarities are worth studying in their own right. Beyond that, there are both real similarities and real differences in the major religions that trace back to their foundations. It is careless and inaccurate to say that all religions teach the same. It is also unfounded to fear that all discussion of differences must be subjective and partisan. There are objective, realistic ways to study the differences and to understand the different focus you find in each faith.

Monday, February 04, 2013

Picture this: What some "alternative" gospels have in common

As basic as my Venn diagramming skills are, I will not try to represent more than four things at a time quite yet. There are more "alternative" gospels than these four, but I have limited myself to these for right now. Here are the reasons why I started with these:
  • The Gospel of Mary, because it has received some publicity and more people may be aware of it
  • The Gospel of Thomas (Coptic), because it is the most focused on Jesus
  • The Gospel of Philip, because it is substantially longer than many alternative gospels (some of which are only a few pages long)
  • The Gospel of Truth, because it is among the earliest-written of the ones that did not make the Bible


Points of interest

To compare this diagram to the previous one, we can immediately see that these alternative gospels are not as tightly related to each other as the Biblical gospels. We see that in an objective, measurable way: they do not all share a core set of keywords, or share the same most common word.

With a high-level picture like this, a picture that shows only the ten most-common keywords from each document, some of the documents don't share any keywords at all. The Gospel of Philip doesn't have an overlap with the Gospel of Mary at that level, and neither does the Gospel of Truth. The Gospel of Mary only relates to the Gospel of Thomas through a match between "Savior" in the Gospel of Mary and "Jesus" in the Gospel of Thomas (more on that in the technical notes). In general, these documents have as much that is different from each other -- sometimes much more that is different -- than they have in common.

Again, comparing the alternative gospels to the Biblical gospels, "Jesus" does not have as much of a place here as in the Biblical accounts. In some alternative gospels, "Jesus" (or a title like "Savior" to represent him) isn't one of the top ten keywords. For example, in the Gospel of Philip and the Gospel of Truth, "Jesus" is not represented in the top ten words, being relatively less important in those documents than other keywords or concepts. This is in contrast to the Biblical gospels where "Jesus" is the highest-frequency keyword -- the topic of first importance -- in all four.

It is questionable whether "Gospel" is an accurate thing to call those particular documents that make no attempt to relate the life or teachings of Jesus, and where "Jesus" does not appear prominently in the top keywords. If a gospel is something that has -- or claims to have -- some record of the life and teachings of Jesus from a viewpoint of his early followers, then some of these documents simply don't meet that definition. Neither is it any complaint against certain documents to mention they do not meet that definition; they simply were not written with the purpose of recording the life and teachings of Jesus. (I've mentioned before, "unorthodox patristics" may be a more accurate classification for some of these documents.) Bear in mind that we are measuring the emphasis of all the documents in an objective way; the result is not something that depends on any ideology, belief, or lack of belief. We are not measuring things one way for the Biblical documents and another way for the non-Biblical documents: we have a level playing field, and all documents are handled in the same way. When we handle the documents the same way, we find there are some differences in what they contain. We are letting each document's contents set its own keywords, and measuring how often those keywords are used. Anyone would get the same results regardless of their views on the documents in question. It is not some sort of ideological thing to notice that some documents are mainly about Jesus while others are not; it is a matter of objective fact. Calling these other documents "gospels" may generate publicity, but it is done at the expense of accuracy about what they actually contain.



Technical Notes

  1. In the section where the Gospel of Mary and Gospel of Thomas overlap, "Savior" from the Gospel of Mary is treated as a match to "Jesus" in the Gospel of Thomas. The Gospel of Mary never actually mentions the name of the "Savior" in the remaining text that we have. Matching "Savior" to "Jesus" is a debatable move in a word matching exercise -- not because the identity of the "Savior" being discussed is in serious doubt, but because "Savior" is a religious idea or title, while "Jesus" is the name of a historical person. These differences reflect something you see when you compare the Gospel of Thomas and the Gospel of Mary side-by-side. The Gospel of Thomas aims to record the sayings of Jesus; it consists largely of a series of sayings introduced by the phrase "Jesus said". In the Gospel of Mary, the Savior makes only a brief appearance in person (at least in the remaining text that we have), and he is more often the topic of discussion among Mary and the disciples. So there is some question about allowing a word match between "Savior" and "Jesus" since they aren't actually the same word and are used somewhat differently. However, if we did not allow "Savior" and "Jesus" as a match in this case, it would give the impression that the Gospel of Mary wasn't related to the others at all; I thought it would be a more accurate representation to show that there was a match of sorts.
  2. There is a second difference between how these charts were made compared to the earlier one. Previously, when we looked at the Biblical gospels, all four shared the same most-frequently-used keyword: Jesus. To reflect this, "Jesus" was in larger typeface, and all the keywords were left in black letters since there was no need to distinguish a different keyword for the different documents. Here, the four writings have four different most-frequently used keywords. For Mary it is "Savior"; for Thomas it is "Jesus"; for Philip it is "man", and for Truth is it "father". To represent this, I have put the most-frequently-used word in the same color as the circle to which it belongs (that is: red for the most-frequently-used word for the Gospel of Thomas, and so forth).
  3. For a different type of technical note: I have been working on my diagramming skills a little bit, and I hope the next diagram will be slightly more detailed without losing its readability.

Thursday, January 31, 2013

Picture this: what the Biblical gospels have in common

As I begin to show the point of the recent series with all the data analysis, I'd like to translate some of those endless statistics into pictures. That should make the point more visible and easier to understand without having to refer back to charts full of statistics.

Because my Venn diagram skills are basic, I kept this chart basic as well: it shows the ten most-used keywords in each of the Biblical gospels. (It's tempting to try to pack in more information. The circles could change size to reflect the size of the original document, or the words might change size or color. Tempting, but for the moment it's overkill.)

This picture gives you a high-level overview of the four Biblical gospels. They have a lot in common, and all four of them share the same most-used keyword: Jesus. They also each have areas that are uniquely their own.

As I go forward, I need to add a few more word clouds to the ones available here on the blog, and then I intend to put together some additional pictures to illustrate what you really see when you compare certain things objectively.

Sunday, January 27, 2013

The Gospels, the Tao, and the Analects: Comparison

There are many reasons we might want to compare two documents to see how much they cover the same material. We have looked at the Biblical gospels in comparison to each other, and to one of Paul's letters. We have compared the combined gospels to the Torah. We have looked at how the Biblical gospels compare to a Gnostic Gospel. Here we take it to the next step: What do we see when we compare the gospels to the texts of other religions?

While I eventually want to analyze far more texts than these, I started by comparing two Biblical gospels (Mark and John) to two eastern texts (the Tao Te Ching and the Analects of Confucius). Full disclosure: I'm fond of both the Tao and the Analects, and am starting here because I am glad for a chance to re-read them and review them again. I considered writing up the comparisons separately for the Analects and the Tao, but there is more that comes to light when the comparisons are reviewed side-by-side.

Summary of Results 

First, comparisons of two Gospels and the Tao
Gospel of Mark and the Tao: 7% shared emphasis (or 10% if "teachers" and "sages" are matched)
Gospel of John and the Tao: 9%

Next, comparisons of two Gospels and the Analects
Gospel of Mark and the Analects: 22% shared emphasis
Gospel of John and the Analects: 18% shared emphasis

For those interested, comparison of the Tao with the Analects:
16% shared emphasis, or 20% if "Master" and "sages" are matched.



Details

The Gospels and the Tao have so low a match that it barely registers. The match between Mark and the Tao is the result of only 5/48 words from Mark's keywords list: people, teachers, things, called, and heaven. Again, the match between John and the Tao is the result of only 5/44 words from John's keywords list: world, life, things, people, called. We may know that both are on the general topic of teaching people about life, the world, and heaven -- a very high-level, summarized type of common ground.

The Analects, on the other hand, have a noticeably higher match. For Mark and the Analects, there are 9/48 words matched: man, asked, people, replied, things, heard, called, heaven, others. For John and the Analects, there are 9/44 words matched: man, asked, love, replied, heard, things, people, speak, called.

So the reason the Analects is more similar to the Biblical gospels is mainly from the basic framework of the documents: the Analects, like the gospels, narrate someone's teachings through their conversations with others. I would wonder whether there would be a similar patten found for any writings that record dialogue-style conversations, especially teachings.

For "called", it should be mentioned that a word may have more than one meaning, and a next-generation version of this tool would eventually need to take that into account. A disciple may be "called" by Jesus, and an act may be "called" virtuous, without "called" really meaning the same thing. That is to say, this version of the tool may slightly miss its estimate since it does not have that kind of precision yet.

For a little more perspective, when we compare the Tao to the Analects, we find 8/50 keywords matched: people, virtue, called, things, heaven, wish, state, words. If we consider "sages" and "Master" as a match -- which is debatable -- that would be 9/50 keywords matched.

The Tao and Analects share some things with each other that they do not share with Mark or John. To take one example, they share an emphasis on "virtue". Anyone who has read the gospels knows that human "virtue" is found under different words; it is not necessarily easy to say which is the closest match. Do we compare the call to be "righteous" or "perfect" or "holy"? Or do we note that the gospels take a different approach from discussing the hypothetical man of virtue? These are not questions I will pretend to answer in a mathematical analysis of word frequencies. There are some kinds of questions that the mathematical analysis may answer; for others, it simply brings to our attention other areas that deserve a look.

The Tao and the Analects were written in different languages than the Biblical gospels; the comparisons have all been done from English translations. (Beyond that, they also came from different cultures and were speaking to different contexts.) I don't apologize for comparing them in English since eventually we have to find a common platform on which to compare them. Part of the job will be to keep that common platform from distorting the picture, no matter which common platform is chosen. The original languages and cultures will need to remain part of the picture.


Sunday, January 20, 2013

The Biblical Gospels and a Gnostic Gospel: Comparison

I have been exploring what you can learn from an objective, mathematical analysis of documents like the gospels to get an idea of their core subject matter, and how that compares to other documents. For the next step, I would like to compare two Biblical gospels to a Gnostic gospel. (Given time, I'd like to compare each of the Biblical gospels to each of the alternative gospels, but I have to start somewhere.)

Summary of Results

Gospel of Mark and Gospel of Truth: 10% shared emphasis match
Gospel of John and Gospel of Truth: 22% shared emphasis match if "sin" and "deficiency" are considered different things; that would be a 23% match if "sin" and "deficiency" were considered matching.

The less-precise estimates, the "Shared Word Estimates", were 7/48 for the Gospel of Mark and Gospel of Truth, and 12/44 for the Gospel of John and the Gospel of Truth (or 13/44, if we consider "sin" and "deficiency" a match). I'm curious whether there is a bigger gap between the two kinds of estimates in some circumstances, though the methods I'm developing are still a little bit new for me to have a real feel for the differences there.

Details

I chose two different Biblical gospels for the comparison to a Gnostic gospel because I was fairly sure that would highlight some features of the documents. The Gospel of Mark sticks more closely to telling events, while the Gospel of John reflects more on the perceived meanings, while still narrating some events. The Gospel of Truth does not narrate Jesus' life; in fact "Jesus" does not show up in the key words list, while it is first in all four of the Biblical gospels. But the Gospel of Truth is reflective in nature, pondering over the perceived meaning of things. These differences show up in that the Gospel of Truth is more similar to the Gospel of John than the Gospel of Mark. (It is still less similar than, say, the Torah is to the combined gospels.) I'd remind the reader that the mathematical tools are fairly new and are still being calibrated; we'd have to look at a good number of documents to see what normal ranges might be and get a clearer idea how to interpret the numbers.

For the Gospel of Mark, there are only 7 words contributing to the match to the Gospel of Truth, words that are high frequency in both documents: son, spirit, gave, things, called, father, truth. These are fairly generic words (for example, "thing" could be anything), and none of the matches is as much as 3% of the high-frequency words in Mark's gospel.

For the Gospel of John, there are 12 words contributing to the match to the Gospel of Truth: father, son, truth, comes, things, light, spirit, whom, gave, himself, speak, called. While the Gospel of John's top word, "Jesus", is not important in the Gospel of Truth, the second-place word from the Gospel of John, "father", is the first-place word in the Gospel of Truth. The three words "father", "son", and "truth" account for roughly half of the "matched emphasis" of these documents. There is a question whether "sin" and "deficiency" should be considered a match; it makes roughly a 1% difference in the match rate.

Moving Forward

I suspect this kind of study -- where I select certain documents and run comparisons -- is more interesting to me than the reader, but I have a few more comparisons in mind before I'll be able to make the point clearer. If you'll bear with me patiently, I think the end result will be more generally useful. I'm hoping that, shortly, this post will be like any long and complex math problem where you "show your work". That is, the work itself is mostly shown so that, when you get to the ending conclusions where people ask themselves, "Is that right?", everybody interested can see exactly how you got those results, and check it for themselves. It is necessary to show the work, and have a completely objective and transparent method -- especially if the objective results are not always in line with conventional wisdom.

Friday, January 11, 2013

Beyond the New Testament: Comparing the Biblical Gospels to the Torah

For our next step, I compared the gospels to the Torah. That is, I compared the combined texts of Matthew, Mark, Luke, and John to the combined texts of Genesis, Exodus, Leviticus, Numbers, and Deuteronomy.

The Short Version of the Results

Shared Word Estimate (12/46) = 26% or (11/46) = 24%, depending on whether or not we recognize a match between the word "Israelites" in the Torah and the word "Jews" in the New Testament. We're working with differences in language, history, and culture; you could argue that decision either way. So the results are presented both ways.

Shared Emphasis Estimate: 25% or 24%, again depending on whether "Israelites" and "Jews" are considered a match.

Different than Comparisons within the New Testament

Between the Torah and the Biblical Gospels, we have a match on emphasis that is roughly 24-25%. It is only slightly less than the match rate between Paul's letter to the Romans and the Gospel of Luke, though again far less than matches among the gospels.

Many of the differences are so expected that I will mention them without much comment. While the gospels discuss Jesus and his disciples, the Torah discusses the patriarchs and Moses. The context is different: the gospels have "Jerusalem" or the "house", while the Torah has "Egypt" and "tent". The language of ancient ritual sacrifice has a set of words that are common in the Torah but not in the gospels: offering, altar, sin, blood, fire, holy, gold, burnt, grain, animal. ("Sin" is at the low end of common words in the individual gospels of Matthew and John, but does not make the common words list of the combined gospels. The words specific to the ancient sacrifices are not common in any of the gospels.)

When we take a closer look at the areas that are in common, we see most of the shared emphasis coming from the words "God" and "Lord", "father" and "son", "man" and "people" -- and "priests". Most of those matches call for a closer look and further thought.

If "Lord" generally means "God" in the Torah, but sometimes "God" and sometimes "Jesus" in the gospels, then do those words mean the same thing and should those words really match? A Christian might say yes, a Jew might say no, and both might agree that this difference is critical. It is important not to take the mathematical studies in a way that depends on someone's personal viewpoint; it defeats the purpose of a mathematical review. Someone who is an analyst first might note that both documents have an emphasis on "Lord" as an important figure, while allowing that the different authors may have different ideas about who exactly the Lord is. And that should be a fair observation from anyone's point of view.

If "father" and "son" are used in numerous accounts of family and genealogies in the Torah, does it really count that they are also common words in the gospels, where the words might mean God and Jesus? Then again, the gospels also contain two Jewish-style genealogies. Never mind for the moment whether they match each other or whether you believe either of them; the same could be said of other genealogies. The point for a computerized word comparison is not whether you can post the family tree on a genealogy site; the point is that some of the structure of the gospels of Matthew and Luke -- the structure that includes a genealogy -- is part of a convention that goes back to the Torah. So the commonness of "father" and "son" language in the gospels is not entirely foreign to the Torah, and has some of its roots in the Torah. God is also referred to as "father" in the Torah (see Deuteronomy 32:6, though I have certainly not yet made a complete check to find all the references.) So the references to God as "father" in the gospels are not completely unheard of in the Torah, and may in some part trace back to a tradition found in the Torah. For matches like this, as they say, the truth is complicated.

Moving Forward

This comparison also brings to light some more features of the tools being used. The more different two documents are, the more questions that arise about even the matches that are found. Here the gospels are written in a culture that was deliberately trying to live according to the Torah and pattern their religious thoughts after the Torah, so the gospels and the Torah were bound to have some similarities. The similarity of the gospels to the Torah was roughly on the same scale as a similarity of Luke to the letter to the Romans, though for different reasons.

As this series goes on with a few more examples, we will likely see a few comparisons that show only a slight relationship if any. When documents are largely different, there is another question that comes to light: are the documents similar enough to warrant a comparison? I won't presume to answer that question so soon, with only a few comparisons completed so far. It may be useful to look at two completely unrelated documents at some point to see if the analysis can detect that that lack of relationship.

Wednesday, January 09, 2013

Beyond the Biblical gospels: comparing Luke and Romans using mathematical models

Thanks to all for your patience while I enjoyed the holidays with lots of family time and remarkably little blogging or research. :)

In this post we take our next step with the mathematical models, and it begins to show different kinds of results. To this point we have been looking at the Bible's four gospels: Matthew, Mark, Luke, and John. To expand our horizons a little, the next document I'd like to consider is Paul's letter to the Romans. It is an early letter within the Christian church, it has been vital in the formation of Protestant Christianity. In modern times the question has become more pointed: Did Paul stay with the direction laid out by Jesus, or was Paul responsible for a change of course? I will not presume to answer that question here, but I will point out some promising pieces of objective information that come to light with this kind of mathematical review.

To compare this letter to a gospel, then, I chose the Gospel of Luke. Since Luke was a companion of Paul's, I thought it could be a productive place to begin.

The Short Version of the Results

Shared Word Estimate (13/52) = 25%
Shared Emphasis Estimate: 27%

Much Different than Gospel-to-Gospel Comparisons

For the first time, all of the matching methods show less than a 50% match -- and here the match is significantly less than 50%. While the gospels consistently had a shared emphasis estimate higher than 50%, Paul's letter to the Romans matches Luke at roughly half that level.

There are several kinds of differences that are immediately seen. We will start at the top of the list with the most common word. The gospels all had the same word as the most common word: Jesus. The letter to the Romans has a different most-common word: God. In fact, "Jesus" doesn't appear until #8 on the list in Romans. However, "Christ" appears higher on the list than "Jesus".

What do we make of the fact that "Jesus" is the common way to speak of Jesus in the gospels, but "Christ" is more common in Paul's letter to the Romans? The word "Christ" does not appear on the common-words list of any of the four Biblical gospels. To be sure, even if the word "Christ" is not prominent in the gospels, still the idea that Jesus is the Christ is well-known from the gospels. They all make a point to explain that Jesus is the Christ, and to demonstrate it. In the gospels, the time when Peter identifies Jesus as the Christ is portrayed as a key teaching, and so is the moment at Jesus' trial where the political leaders ask whether Jesus is the Christ. The Gospel of John even explains that the reason the book was written is "that you may believe that Jesus is the Christ". So the concept of "Christ" is an important idea in the gospels, even though the word is not used often. We could say that calling Jesus by the title "Christ" shows the next stage of logical development after those gospel accounts. That is, calling Jesus "Christ" shows a prior acceptance of those teachings about Jesus. It is, in a way, summary-level talk, to call Jesus by the title of Christ. The gospels are interested in explaining and demonstrating that Jesus is the Christ; for the epistle to the Romans, this has already been explained to the readers' satisfaction and is now part of the foundation on which they build. So here we have a new kind of difference: a difference about the level of detail being used or the logical progression of ideas, whether something is demonstrated or already given. It is a difference in the level of the conversation, and in the starting point of the discussion.

But to what extent is it discussing the same subject matter? The action from the gospels, the physical settings and the people who first heard Jesus are not a large part of the picture in Paul's letter. The book of Romans does not commonly speak of "crowds" and "disciples", or "Peter" and "Mary", or "Jerusalem" and the "house", or "asked" and "answered" in the way that the Gospel of Luke commonly does. The actions from Jesus' life are not being narrated in his letter; the letter is a different type of material. Paul does have some interaction with people in his letter, but he interacts with the people that he expects to read his letter. So while there is no "crowd" in Paul, instead we have Paul's trademark where "greet" is on the common word list in Romans, and there is a small crowd reading the letter. (Anyone who reads a few of Paul's letters will notice that he spends a certain amount of time on personal greetings. We know many early Christians by name because Paul greeted them by name in his letters.)

Still, the differences go deeper. The gospels are all biographies, or we might say the fourth gospel is a memoir and reflection on Jesus' life. As records of Jesus' life, all four gospels share the same most common word: "Jesus". The letter to the Romans, on the other hand, has "God" as the most common word, then "sin" and "law". To be sure, "sin" and "law" are discussed in the gospels -- but not always enough to make the most-common-words list. For "sin" we may remember conversations about sins being forgiven. For "law", there are records of discussions between Jesus and other people over the interpretation of the law. Questions come up about matters of divorce, or tax, or ritual hand-washing, or which are the most important commandments, or a case of capital punishment, or whether certain religious leaders could claim that Jesus was morally in the wrong for performing miracles to heal people on the Sabbath, as it was a kind of work. So "sin" and "law" both have a presence in the gospels, either directly or by example. Paul discusses these ideas at a summary level, where "sin" and "law" are often abstractions. The same might be said of "faith" and "grace". These words are commonly used in Romans where Paul discusses them in a relatively abstract way. In the gospels these same words "faith" and "grace" are not often used directly, but are instead shown in living action.

But the major differences are not limited to the fact that Paul is more abstract, while the gospels show Jesus in action. By Paul's leading words in Romans ("God", "law", "sin"), we see Paul also trying to put Jesus in a context that his readers might know. He explains Jesus against a background familiar to his fellow Jews, back in his day when the Temple still stood in Jerusalem and sacrifices were still offered daily, where people made pilgrimages for the Torah's decreed feasts, where Torah-based Jewish legal courts had some degree of legal authority and might have jurisdiction over some cases, where someone might comment publicly about a lack of morals if someone failed to perform a ritual washing before a meal, where breaking the Sabbath might lead to a formal legal inquiry. We see Paul struggling with the question: For a Jew like him or many of his readers -- learning that Jesus is the Messiah and that the Messiah is about God's love, about grace and mercy, about good news and life -- what does that mean for their old understanding of law and sin? What does that mean for their ideas about righteousness before God?

We also see Paul spending some effort discussing "Jews" and "Gentiles", "Israel" and being "circumcised". What does it mean that even Gentiles are now included in a new covenant with God? What does it mean that Gentiles have a righteousness before God that did not come from the Law of Moses? What does that mean for whether the Law of Moses should apply to Gentiles? On the one hand, if the Gentiles do not need to be circumcised to be in the New Covenant, then is circumcision still relevant? On the other hand, if Gentiles are now numbered among God's chosen people -- which previously had meant Israel -- then is there still any advantage in being a Jew? Paul considers the implication that God has made a covenant for all people through Jesus; and Paul seems to have something of an identity crisis on what it means to be Jewish now, in light of God opening the gate wide to all nations. For him and his concept of his beloved Jewish nation's role in the world, it is not an easy transition to go from being an only child to being firstborn among many. This emphasis raises the question: to what extent was the letter to the Romans about universal themes for all people of all times, and to what extent was that letter meant to speak to the existential crisis of Judaism that Paul saw in God's new covenant for all nations? (In a few places Paul seems defensive of the special role of his people, and mentions several times that things are "first for the Jew" and then for the Gentile. I should mention that Paul's letter to the Romans is not the only philo-Semitic writing in the New Testament. I have become curious whether anyone has actually studied the philo-Semitism of the New Testament. In my readings through the materials, philo-Semitism seems far more prominent than any supposed "anti-Semitism", which is not surprising since most of the authors were themselves Jewish.)

A few advantages of the mathematical comparisons
  • We may be able to determine whether something is rightly a classified as a "gospel" by whether it is mainly focused on Jesus. It may also matter whether the action/narration words, setting, and character names are still in a prominent place. 
  • We may be able to tell that a document is "next generation" material (from a logical point of view) if it starts by assuming that Jesus is the Christ, as shown by a high usage of the word "Christ" compared to "Jesus".
  • We may need to look for relationships between key words -- like between "Jesus" and "Christ" -- where the difference shows that a historically earlier viewpoint is now taken as "given".
  • To interpret the findings correctly, we may need to look for detail v. summary types of differences, or specifics compared to abstractions, like Jesus' kind encounters with various people as opposed to Paul's mention of "grace" or "mercy".
  • The details of the differences between two documents can show, objectively, where the focus of an author lies and bring out themes that might be missed otherwise.

Friday, December 28, 2012

Comparing Mark and John with Mathematical Models

Thank you for your patience with these document comparisons. We're getting close to my being able to show you some more interesting things you can see with the comparisons, but wanted to at least get all the Biblical gospels into the mix before we started going beyond them. So for the fourth gospel, here are results from comparing Mark with John.

The short version of the results

I did have a chance to work through the problems with the calculation and to make them more sound, where two shared words will now never have a negative impact on the comparison. I'm now simply using the smaller of the two numbers for any word pair, which works out to 0 when the word isn't on both lists. I'll also be updating the previous documents with the corrected calculations.

Shared Word Estimate (22/48) = 46%
Shared Emphasis Estimate 54%

There is less similarity between Mark and John than we previously saw between Mark and Matthew or Luke. In the notes on the Shared Emphasis Estimate, I'll include some notes on where the differences are found.

Notes on the Shared Word Estimate

Again, Mark is the shorter document. It has 48 words included in the high-frequency word list, which is limited to words that would make at least a 1% difference in the total as discussed previously. Of those 48 words, only 22 are also in John's high-frequency words list calculated in the same way, which is the lowest match rate we have seen yet among the gospels. So 22/48 = 46%, rounded to the nearest whole number. Again, since the percentages involved are already effectively rounded when we leave out low-frequency words, it does not seem warranted to use a lot of decimals in the percentage.

Notes on the Shared Emphasis Estimate

With Mark and John, , the highest-frequency word in both documents is "Jesus". But the differences start as early as the second word on the list, where "man" is second in Mark's but "father" is second in John's. For the first time in our comparisons, even though John is the longer document, its high-emphasis words list is actually shorter at 44 words. This is an objective, verifiable measure of what people have long perceived about the fourth gospel: John's perceptions are more distilled or filtered, more focused -- possibly more edited, or more selective.

When we look at where the differences occur, there are some points of interest. Again, the comparisons is done from the perspective of Mark's gospel; other differences would come to light when using John as the baseline. When comparing Mark's top words to John's, there are 26 that are not on John's top words list; they are listed in the order of their importance in Mark's word list: crowd, teachers, around, anyone, began,  took, house, law, against, kingdom, mother, boat, hands, eat, days, lord, children, heaven, others, sitting, twelve, chief, evil, hear, James, looked. When we follow the leads that are given here, we might find fewer crowd scenes and fewer action scenes in John than in Mark.

Then there are the words on both lists that are emphasized noticeably less in John than in Mark: people and man. Again, this adds weight to the possibility that we'll find measurably fewer crowd scenes and action scenes in John.





The histories passed down about the Gospel of John mention that it was written to supplement the previously-written gospels. One way this may be seen is Mark's relatively greater emphasis on Jesus' public life, and John's relatively greater emphasis on private moments.

Tuesday, December 18, 2012

Comparing Mark and Luke with Mathematical Models

I promise there is a point to these document comparisons. I haven't yet calculated all of the comparisons that I intend, but I have read the documents in question, and I have no doubt that an objective, computer-based comparison like this will turn up interesting results. In the meantime, I did notice a few things when comparing Mark with Luke that might interest the general reader.

The short version of the results

Shared Word Estimate 65%
Shared Emphasis Estimate 64%*
* The originally listed number of 53% had some problems where, for word pairs with large differences, the shared word value might be less than the smaller of the two numbers or even negative. This number should be a more solid reflection of what is shared between the two documents.

There is less similarity between Mark and Luke than we previously saw between Mark and Matthew. In the notes on the Shared Emphasis Estimate, I'll include some notes on where the differences are found.

Notes on the Shared Word Estimate

Mark is a shorter document and has 48 words included in the high-frequency word list, which is limited to words that would make at least a 1% difference in the total as discussed previously. Of those 48 words, 31 are also in Luke's high-frequency words list calculated in the same way. So 31/48 = 65%, rounded to the nearest whole number. Again, since the percentages involved are already effectively rounded when we leave out low-frequency words, it does not seem warranted to use a lot of decimals in the percentage.

Notes on the Shared Emphasis Estimate

Again, the two highest-frequency words are the same between the two documents: "Jesus" and "man". And again Luke's list is broader than Mark's: it contains 52 words in the high-frequency list. When we look at where the differences occur, there are some points of interest.

When comparing Mark's top words to Luke's, there are 17 that are not on Luke's top words list: son, around, anyone, mother, Peter, boat, hands, eat, days, others, sitting, truth, twelve, chief, evil, James, and looked. Then there are the words emphasized noticeably less in Luke than in Mark: Jesus (though still by far the top word) and disciples. Some of the less-used words are related: Peter, twelve, James, and disciples. There seems to be noticeably less emphasis on the disciples in Luke than in Mark. That is consistent with early accounts that Luke was a companion of Paul's, showing less interaction with Jesus' disciples than is found in Mark.

I have noticed one problem with the calculations up to this point: the original calculation can cause two shared words to have a negative net effect, if the difference between the frequencies is larger than the original frequency itself. It may give more accurate results to simply use the smaller of the two frequency scores for the words in question, which may be 0 if the word is not found in the second document. At any rate I will finish up a few more sample comparisons before trying any updates to the calculation.

Thursday, December 13, 2012

Comparing Mark and Matthew with Mathematical Methods

Here is the first analysis of actual documents with the mathematical models discussed previously. I've taken my first document as the Gospel of Mark and the second as the Gospel of Matthew, using the word clouds linked here.

The short version of the results

Shared Word Estimate 77%
Shared Emphasis Estimate 69%*
* The originally listed number of 57% had some problems where, for word pairs with large differences, the shared word value might be less than the smaller of the two numbers or even negative. The recalculated number given above should be a more solid reflection of what is shared between the two documents, as it simply uses the lesser of the two values, which is never lower than 0. 

In the notes on the Shared Emphasis Estimate, I'll mention some other things that the statistical analysis shows: with the breakdown done at this level, you can do more than estimate how much is shared. You can also identify where the differences are.

Notes on the Shared Word Estimate

Mark is a shorter document and has 48 words included in the high-frequency word list, which is limited to words that would make at least a 1% difference in the total as discussed previously. Of those 48 words, 37 are also in Matthew's high-use words list calculated in the same way. So 37/48 = 77%, rounded to the nearest whole number. (Since the percentages involved are already effectively rounded by the exclusion of low-frequency words that would chip away at the percentage, I don't think a lot of decimal points are significant in the analysis.)

Notes on the Shared Emphasis Estimate

When it comes to the detail matching on emphasis, the two highest-frequency words are the same between the two documents: "Jesus" and "man". Matthew's list is broader. It contains 53 words in the high-frequency list. So words are generally lower-frequency in Matthew than they are in Mark. This raises a question about the method, whether some sort of adjustment is in order for the relative length of the lists. It's worth considering, but my first thought is that if we're measuring relative emphasis, and the relative emphasis were the same between documents, then the word frequency lists would be the same between the documents. So my first inclination is not to adjust for different list lengths, but to consider that difference as part of an accurate reflection that the two documents have a somewhat different emphasis.

The emphasis estimate turns out to yield more information than the originally-intended measure of how much two documents are alike. It also gives some insight into what exactly is different. So with that in consideration, the words showing the biggest difference in emphasis are "Jesus" which is emphasized somewhat less in Matthew though it is still by far the most frequent word, then "father" and "heaven" which are used noticeably more in Matthew than in Mark. Those three words account for about 10% points in the emphasis-gap between the documents. Another significant gap comes from the 11 words on Mark's list but not in Matthew's: around, began, boat, hands, days, sitting, twelve, evil, hear, James, looked. That is not to say those words don't occur in Matthew, but that they don't make the high-frequency words list as they do in Mark.

Any areas which show a difference in emphasis might be worth closer study. I find it interesting that such a practical, ordinary word as "boat" should make the high-frequency list of Mark. The early records we have about Mark say that he was writing about Jesus as told to him by one of the disciples who was a fisherman by trade. The relative emphasis on the "boat" in Mark does not prove that the source of information was a fisherman, but it is consistent with that possibility. It might indicate an area for further research, to see what kinds of information might come to light by taking a closer look at the "boat" references in Mark. The "father" and "heaven" emphasis in Matthew over Mark might also bear a closer look. Other differences (like "around" or "began") seem less promising, though it would still be best to do a quick check of the original texts to make sure that it is just a difference in narration style or something of that sort.

Monday, December 10, 2012

Mathematical Methods for Comparing Document Content

This post describes two different ways to calculate a rough measure of how much two documents cover the same material or the same topic. This is mainly written for those who want technical details of how the calculations work; it may not be of interest to other readers. The two methods are a "shared word estimate" and a "shared emphasis estimate".

Two Sample Lists

Imagine two very short books with the following word frequency lists: 

Document #1
  1. fun (15)
  2. Dick (12)
  3. Jane (10)
  4. Spot (6)
  5. see (4)
  6. run (3)

Document #2:
  1. fun (28)
  2. Fred (27)
  3. Jane (25)
  4. Dick (13)
  5. catch (4)
  6. Spot (3)

(I haven't really run the numbers for any actual "Dick and Jane" early-reader books; these numbers are made up for the purposes of illustration.) The two measures that I'll calculate show the percentage of shared words on these lists, and then a more detailed comparison of their emphasis.

Calculating A Shared Word Estimate

For the shared word estimate -- a rough estimate of whether the documents cover similar material or subject matter -- we run a basic count of the words shared between the two lists, and compare that to the length of the list. In this simple example, each list has 6 words, shares 4 words with the other list, and contains 2 words not found on the other list. So the shared word estimate tells us that 4/6 (67%) of the common words are the same between the two lists. The shared word estimate is crude, but can be used as a first estimate of whether a more detailed comparison is in order. You can determine, mathematically or by computer analysis, that these two documents may be related. If you saw a book with another top words word list, like "eggs, green, ham, am, Sam, like", you would find 0% in common and could expect that this document was not covering the same material or narrative.

A quick look at the shared word estimate shows that there is room for improvement, though. If the top, most common word is the same on both lists, there is a higher chance that they are on the same topic than if the bottom words happen to match. A more detailed comparison is in order that takes things like that into account.

Calculating A Shared Emphasis Estimate

The "shared emphasis estimate" measures not only whether both documents use the same words commonly, but considers whether those words occur about as commonly: it measures emphasis as well. Here the first approach I tried based on word rank (how high a word scores on the list) had to be discarded, as there were significant problems with the validity of the result. Simply comparing the rank of each word from one list to the next did not account for the fact that some lists have near-ties at some places, while others have steep drop-offs in word frequency, meaning that the ranking number was not an especially clean measure of the commonness of a word. The longer the list, the greater the problem that would be presented. Then there was a question of how much to weight the first-ranked word compared to the second-ranked, and so on down the list. If we used the rank as a basis for the weight, it would introduce an inflexible and artificial scale. The more fitting method is to weight each word based on its prevalence within the documents in question.

To determine the weight for each word, then, first a total was run of all the word-occurrences in the list. Then each individual word's usage count was turned into a percentage of that total. Here are our two sample documents again, with those calculations shown:

Document #1: 50 total count for the words in the "common words list":
  1. fun (15): 15/50 = 30%
  2. Dick (12): 12/50 = 24%
  3. Jane (10): 10/50 = 20%
  4. Spot (6): 6/50 = 12%
  5. see (4): 4/50 = 8%
  6. run (3): 3/50 = 6%
Document #2: 100 total count for the words in the "common words list":
  1. fun (28): 28/100 = 28%
  2. Fred (27): 27/100 = 27%
  3. Jane (25): 25/100 = 25%
  4. Dick (13): 13/100 = 13%
  5. catch (4): 4/100 = 4%
  6. Spot (3): 3/100 = 3%
To calculate the Shared Emphasis Estimate, we take each word's emphasis percentage in the first document as our starting point. Comparing it to the second document, we subtract out the difference in how much it is emphasized there to find the shared emphasis between the documents.

For example, "fun" has 30%  value in the first document, but 28% in the second. The difference in emphasis is 2%. So the shared emphasis is 30% - 2%, or 28%, based on the first word. The other words are also added into the result.

A slight miscalculation: This calculation had to be refined because of problems in whether it was actually measuring what was intended. Originally the calculation used the absolute value of the difference, then adjusted the original amount by that, using the calculation below. Here it shows the calculation for each of the six words listed for Document1 compared to Document2, and uses "abs()" rather than "||" to mean absolute value:
  1. fun: 30 - abs(30-28), or 30 - 2, = 28.
  2. Dick: 24 - abs(24-13), or 24 - 11, = 13.
  3. Jane: 20 - abs(20-25), or 20 - 5, = 15.
  4. Spot: 12 - abs(12-3), or 12 - 9, = 3.
  5. see: 8 - abs(8-0), or 8 - 8, = 0. (Not a shared word.)
  6. run: 6 - abs(6-0), or 6 - 6, = 0. (Not a shared word.)
Totaling those numbers, we get 28+13+15+3+0+0 = 59% shared emphasis estimate. The Shared Emphasis Estimate typically will be lower than the cruder Shared Words Estimate. This is because the Shared Words Estimate would give full weight to a match between the least-used word and the most-used word, and takes no account of differences in emphasis.

Updated calculation: The original calculation worked acceptably well for the simple and hand-made examples above, but when comparing actual documents some problems appeared. Consider percentages like the following for a pair of words:

List1: 2%, List2: 15%.  

The absolute value of the difference is 13%, and subtracting 13% from 2% we get -11%. It is then possible for a pair of words to have a negative impact, even when it appears in a significant way in both documents. Based on what I am intending to measure, the number that should be used is simply 2%, the smaller of the two numbers.

Or consider the following example:

List1: 2%, List2: 3%.  

The absolute value of the difference is 1%, and subtracting 1% from 2% we get 1%. But each document has at least 2% value for that word, so it is a more accurate reflection of what I'm intending to measure if the shared value is 2%.

The refinement to the calculation is to leave out the absolute value of the difference, and simply take the smaller of the two numbers for any given pair. This will be zero when the word is on one list but not the other, but it will never be less than zero. 
  1. fun: lesser of 30 or 28: 28.
  2. Dick: lesser of 24 or 13: 13.
  3. Jane: lesser of 20 or 25: 20.
  4. Spot: lesser of 12 or 3: 3.
  5. see: lesser of 8 or 0: 0 (Not a shared word.)
  6. run: lesser of 6 or 0: 0 (Not a shared word.)
Figuring the totals again: 28 + 13 + 20 + 3 + 0 + 0 = 64% for the shared emphasis estimate. For most of the pairs the result was the same, but now the "shared emphasis" is never less than the smaller of the two amounts, which is a more accurate measure of what that calculation is intended to show.

Further Refinements

Here I worked with two very basic (and fictitious) sample documents, where I had the prerogative of selecting the values used for the example. In real documents, another question is significant: how many words do we compare? Here we compared six words, but that was arbitrary. What is a sound method for determining how many words to include in the comparison?

Since this method is generally intended for longer works, my starting point is this: each word is added to the list in order of decreasing usage, with the most-used word being added first, followed by the second most-used word, and so forth. (Some structural words such as articles and conjunctions are typically filtered out during word counts.) When adding each new word to the word list, keep going so long as the current new word, if included, would have a value of 1% or more of the total. Once the next word would be less than 1% of the total, that's probably the point at which the additional comparison doesn't refine the result enough to be relevant. In this way, the number of words included in a list is not an arbitrary number, but is sensitive enough to respond to the different word usage characteristics of each document. At the same time, the measure remains objective to the point where the calculation could be done, content-blind, by a computer program.

I worked out the methods for calculating a Shared Emphasis Estimate by using hypothetical sample books until I had a method with an objective basis (one that could be turned into a computer program that is indifferent to the content, given the time to write the code), and that gave reasonable results.

A few potential design problems may need work. First, the 1% rule is for documents of substantial length; it is possible that it would need some amendment for shorter documents such as our mini-documents used as test cases above. I have not yet tried to compare shorter works, but the lists above suggest the problem could be real and, on a short enough document, some rarely-used words would be included simply because the word totals never reached 200, which is the tipping point for excluding words used only twice. It's possible that more of a "bell curve" approach might eventually replace the 1% rule, as something more easily scalable to different sizes of document.

Also, in larger documents especially, there may be ties in how frequently words are used: that is, more than one word might be used at the same frequency. This is common enough in larger documents, and it can happen right at the 1% boundary. In such a cluster of words of the same frequency, it is possible that the first would meet the 1% rule but the last would not if we had already added in that previous word of the same frequency. In that case, it makes no sense to show a preference for one word over another when both have the same frequency. That is to say, if the first word of a certain frequency is included under the 1% rule, then all other words of the same frequency would be included on the list because of their frequency, even if the resulting final percentage for those words might be slightly under 1% when the whole group of words is included.

A future area for exploration would be: how much can we tell about a document's content from this kind of analysis? For example, would a biography typically have the subject's name at the top of the word-frequency list? I would also be curious how different types of political and persuasive material would look, and what kinds of emphasis became apparent. I'd also see some potential for targeted word frequencies: for example, words that frequently appeared only in one portion of a document, or throughout a document but only while discussing only one recurring topic.

Forward


Next we will see how the basic approach works with actual documents instead of hypothetical ones. But that will wait for another post.


Tuesday, December 04, 2012

Can you measure how much are two documents alike?


In my day job as a programmer, I spend a certain amount of time analyzing data, and in the bigger projects there can be millions of records and over a billion individual fields being handled. And each individual field has to be handled correctly by specialized programming routines; designing and testing those is my job. What does that have to do with this blog? Habits carry over from one place to another, and at times I view documents -- for example the gospels, or systematic theology -- as another job in high-volume data analysis. (I know, some people think that sounds really dull. Regardless, it leads to fascinating places.)


I've done a number of word clouds on this blog. They are one way to do a quick, high-level overview of a document. The next question on my mind is: can you get an idea of how closely two documents cover the same material by comparing their word clouds? When you look at a word cloud, you see a graph of the important words for a document. The information used to create that chart is a list of words and a count of how often they appear. I've been looking at ways to take two lists for two documents and estimate how closely those two documents cover the same material. After a few tries that left much to be desired, I have a method which is promising and objective, with the important decisions being based on mathematical criteria rather than human judgment.

What could you gain with a comparison like that? You could get a rough answer to a question like, "How closely does the Gospel of Matthew cover the same material as the Gospel of Mark?" Or "How closely does the Gospel of John cover the same material as the Gospel of Matthew?" How about comparing Paul's letters to the gospels to see how closely they track each other? How about comparing the "alternative" gospels to the Bible's gospels? How about comparing a catechism or some writer's systematic theology to the gospels, or the New Testament, or the Bible as a whole? How about comparing the holy books of one religion to another, to get a feel for similarities and differences?

In upcoming posts I'm hoping to start exploring some of those questions and their answers. In a future post I will also give the mathematics and logic of how the comparison is done, for those interested. Below is the other major point for a general reader: the most important limits of the method.

Limits of the method

The first limit of the method comes from the fact that it is based on word counts: the content is summed up at the word level, without the phrases or thoughts or the relationships connecting them, without any sense of intent or purpose, logic or history. It would be possible for two authors to take very different approaches to the same concepts, and this particular method could not tell the difference if the authors used the same words at roughly the same frequency.

A second limit is the issue of synonyms and near-synonyms. Do we compare "elected" and "chosen" as the same? How about "predestined" and "foreordained"? "Walked" and "went"? Future development would include a way to factor in a weighted, partial match for near-synonyms or similar words during the matching process.

Another limit is the difficulty comparing documents at two different levels of detail. If one document discussed "oaks" and "pines" and "elms" and "maples", and another discussed "forests", this method would not see the "forests" for all the specific trees in the first document. The more different the level of detail, the more noticeable the problem becomes. For example, "Five teenagers get Saturday detention" might be a recognizable reference to the movie The Breakfast Club, but I seriously doubt that a word-cloud comparison of that phrase to the script would identify that they were talking about the same thing. A more fully-developed method would take into account how to move from the specific to the general, and what kind of detail would be the right match as you "zoom out" to higher and higher summary levels.

The method is also limited to checking for one particular type of relationship between documents: it shows documents that are probably covering the same general material. It does not cover other relationships, for example "prequel" and "sequel", "original narrative" and "commentary", or other types of relationships.

It's likely enough that more shortcomings will show themselves as we work through a few examples. But for all the limitations, it should still be a useful estimate of how much two documents cover the same topics.