Showing posts with label dna. Show all posts
Showing posts with label dna. Show all posts

Friday, March 14, 2014

The MAGICAL Number Seven (plus or minus two)

Why do we “chunk” things in groups of about seven – seven days of the week, seven seas, seven sins, etc? The presentation I gave to the Philosophy Club in The Villages, FL, 14 March 2014 provides the theoretical answer. You may download a PowerPoint Show that should run on any Windows computer here: https://sites.google.com/site/iraclass/my-forms/PhiloMAGICALsevenMar2014.ppsx?attredirects=0&d=1

This is an easy-to-understand version of a more technical presentation I made to the Science-Technology Club in February, see: http://tvpclub.blogspot.com/2014/02/optimal-span-amazing-intersection-of.html


MILLER - PERSECUTED BY THE NUMBER SEVEN !



Way back in 1956 a classic scientific paper appeared in the Psychological Review with the intriguing title: The Magical Number Seven, Plus or Minus Two – Some Limits on Our Capacity for Processing Information. That paper was extremely important and influential and is still available online. George A. Miller begins with a strange plea:
My problem is that I have been persecuted by an integer … The persistence with which this number plagues me is far more than a random accident …
He presents the results of twenty experiments where human subjects were tested to determine what he calls our "Span of Absolute Judgment", that is, how many levels of a given stimulus we can reliably distinguish. Most of the results are in the range of five to nine, but some are as low as three or as high as fifteen. For example, our ears can distinguish five or six tones of pitch or about five levels of loudness. Our eyes can distinguish about nine different positions of a pointer in an interval. Using a vibrator placed on a person's chest, he or she can distinguish about four to seven different level of intensity, location, or duration, etc. The average Span of Absolute Judgment is 6.4 for Miller's twenty one-dimensional stimuli.

Miller also presents data for what he calls our "Span of Immediate Memory", that is, how many randomly presented items we can reliably remember. For example, we can remember about nine binary items, such as a series of "1" and "0", or about eight digits, or about six letters of the alphabet, or about five mono-syllabic words randomly selected out of a set of 1000.

At the end of his paper Miller rambles:
...And finally, what about the magical number seven? What about the seven wonders of the world, the seven seas, the seven deadly sins, the seven daughters of Atlas in the Pleiades, the seven ages of man, the seven notes of the musical scale, and the seven days of the week? What about the seven-point rating scale, the seven categories for absolute judgment, the seven objects in the span of attention, and the seven digits in the span of immediate memory?

For the present, I prefer to withhold judgment.

Perhaps there is something deep and profound behind all these sevens, something just calling out for us to discover it.
But I suspect that it is only a pernicious, Pythagorean coincidence. [my bold]
Well, it turns out that there IS something DEEP and PROFOUND behind "all these sevens" and I (Ira Glickstein) HAVE DISCOVERED IT. And, my insight applies not only to the span of human senses and memory, but also to the span of written language, management span of control, and even to the way the genetic "language of life" in RNA and DNA is organized. Furthermore, my discovery is not simply based on support from empirical evidence from many different domains, but has been mathematically derived from the basic Information Theory equation published in 1948 by Claude Shannon, and the adaptation of "Shannon Entropy" to the Intricacy of a biograph by Smith and Morowitz in 1982.


SIMPLICITY VS COMPLEXITY VS INTRICACY 


Albert Einstein wisely advises us to:
Make things as simple as possible, but no simpler!
Good advice, but how to follow it? Well, Edward Teller suggests:
Threads of simplicity … are not easily discovered in music or in science. Indeed, they usually can be discerned only with effort and training. Yet the underlying simplicity exists and once found makes new and more powerful relationships possible.


How to find those "threads of simplicity"? Well, we need to understand the difference between COMPLEXITY and INTRICACY which, in normal usage, are sometimes used interchangeably. However, there is an important distinction between them according to Smith and Morowitz (1982).

COMPLEXITY - Something is said to be complex if it has a lot of different parts, interacting in different ways. To completely describe a complex system you would have to completely describe each of the different types of parts and then describe the different ways they interact. Therefore, a measure of complexity is how long a description would be required for one person competent in that domain of knowledge to explain it to another.

A great example of UNNECESSARY COMPLEXITY is found in those "Rube Goldberg Inventions" where a relatively simple task is complicated by combining all sorts of different effects into a chain of ridiculous interactions.

INTRICACY - Something is said to be intricate if it has a lot of parts, but they may all be the same or very similar and they may interact in simple ways. To completely describe an intricate system you would only have to describe one or two or a few different parts and then describe the simple ways they interact.

A great example of useful INTRICACY is a window screen that is intricate but not at all complex. It consists of equally-spaced vertical and horizontal wires criss-crossing in a regular pattern in a frame where the spaces are small enough to exclude bugs down to some size. All you need to know is the material and diameter of the wires, the spacing betwen them, and the size of the window frame. Similarly, a field of grass is intricate but not complex.

If you think about it for a moment, it is clear that, given limited resources, they should be deployed in ways that minimize complexity to the extent possible, and maximize intricacy! That is why nearly all natural and artificial structures are configured as Hierarchical and have a common "Optimal Span".


 A SIMPLE FORMULA FOR MAXIMIZING INTRICACY AND REDUCING COMPLEXITY

What is Optimal Span?

With so many different types of hierarchical structures, each with its own purpose and use, you might think there is no common property they share other than their hierarchical nature. You might expect a particular Span of Control that is best for Management Structures in Corporations and a significantly different Span of Containment that is best in Written Language.

If you expected the Optimal Span to be significantly different for each case, you would be wrong!
According to System Science research and Information Theory, there is a single equation that may be used to determine the most beneficial Span. Thatoptimum value maximizes the effectiveness of the resources. A Management Structure should have the Span of Control that makes the best use of the number of employees available. A Written Language Structure should have the Span of Containment that makes the best use of the number of characters (or bits in the case of the Internet) available, and so on.

The simple equation for Optimal Span derived by [ Glickstein, 1996 ] is:

So= 1 + De
(Where D is the degree of the nodes and e is the Natural Number 2.71828459)

In the examples above, where the hierarchical structure may be described as a single-dimensional folded string where each node has two closest neighbors, the degree of the nodes is, D = 2, so the equation reduces to:

So= 1 + De = 1 + 2 x 2.71828459 = 6.43659

“Take home message”: OPTIMAL SPAN, So = ~ 6.4

Also see Quantifying Brooks Mythical Man-Month (Knol) , [Glickstein, 2003 ] and [ http://repository.tudelft.nl/assets/uuid:843020de-2248-468a-bf19-15b4447b5bce/dep_meijer_20061114.pdf] for the applicability of Optimal Span to Management Structures.
Ira Glickstein

Saturday, August 8, 2009

We Need a Comprehensive DNA Database

A recent TV newsmagazine featured the story of a woman who was raped some 20 years ago, before DNA was generally available to confirm the identity of the suspect. She testified that she carefully observed the facial features of her assailant and helped the police sketch artist make an excellent drawing. She then picked Ronald Cotton out of a photo lineup and later the same guy out of a physical lineup. On the basis of her certain eyewitness testimony, Cotton was convicted and sent to jail.

About 15 years later, another inmate, Bobby Poole, was assigned to the same jail. Poole looked so much like Cotton the guards sometimes called them by each other's name. Cotton appealed for DNA tests against the rape kit that had been preserved by the police. The tests proved Cotton did not do the rape. They also proved that Poole did. Poole was convicted and Cotton was released after spending a decade and a half in jail for a crime he did not commit. Cotton graciously forgave his mistaken accuser.

Cases like this show how unreliable eye-witness reports may be, even if (as in this case) the victim was highly intelligent, took care to be observant, and she and the police and the trial court were totally honest and professional.

According to the TV program, several hundred wrongly-convicted inmates have been released in the past decade on the basis of newly available DNA technology. That is a tremendous stride for justice!

HOWEVER RAPES AND OTHER VIOLENT CRIMES STILL OCCUR

While DNA tchnology is now available to confirm the identity of the rapist if, as in most cases, a DNA sample can be obtained, rapes and other violent crimes continue to occur with disturbing frequency.

The problem is that DNA is used only to confirm identity. The police have to use far less certain, old-fashioned methods to track down the suspect. They must depend upon eye-witness evidence that is known to be unreliable. They depend upon informants who are often criminals themselves and may have their private agendas. They depend upon stereotypes and -lets admit it- profiling based on criminal history, race, age, neighborhood, and gender.

WHAT IF A COMPREHENSIVE DNA DATABASE WAS AVAILABLE?

When an automobile is involved in a crime or an accident and the license plate number is caught on video surveillance or is reported by a witness, it is easy to identify the owner of the car and investigate further.

Wouldn't it be great if this was the case with rapes and other violent crimes?

Violent assailants often leave some bodily evidence (ejaculate, hair, saliva, blood, skin, sweat, ...) on the victim and/or at the crime scene. Given a comprehensive DNA database, it would be almost as easy as looking up a license plate number to finger the suspect!

Yes, a careful and thoughtful rapist could wear gloves and a hairnet and use a condom and require his victim to douche, etc., and that would defeat the DNA ID method in some cases. However, most assailants are not that clever.

OBJECTIONS TO A DNA DATABASE

The only rational objection to a DNA database would come from potential rapists and other criminals who don't want to be caught - and their criminal defense lawyers who like a steady income - often paid out of public defender tax dollars.

Yes, there is the issue of "privacy". Many people do not want their DNA (or fingerprints) on file at the FBI or other police agency because they are worried about how such identifying data might be used by a rogue government cracking down on dissidents or other non-favored individuals.

That is not a worry for me. I quite willingly had my fingerprints taken as part of a security check to allow me to work on classified military projects. As far as I know, my fingerprints ar still on file at the FBI.

In any case, for most of us who have a well-documented and fixed place of residence, families, employers, sources of income, bank accounts, credit cards, cars, and so on, we are easily found. Those of us who keep our cell phones on at all times are leaving computerized records of exactly where we have been, minute by minute, every single day. The only people who may benefit from "privacy" are the homeless and jobless, and the criminals who may commit crimes while using YOUR stolen car or cell phone or credit card or identity!

Another issue, more serious, is the possible use of a DNA database to identify individuals who may be susceptable to certain genetic diseases, and the possible use of that information by health insurers to refuse coverage or charge a higher premium. (As a utilitarian, I see nothing wrong with the current actuarial system where young men pay higher auto insurance rates, smokers higher health premiums, people living in wooden houses higher fire insurance, those in tornado alley higher storm insurance, and so on based on demonstrated risk levels. Unfortunately, health insurance seems to be moving into a different category even for illnesses that are mostly self-inflicted due to smoking, drinking, or over-eating.)

The genetic ID objection may be dismissed easily. DNA has sufficient markers such that those associated with genetic deseases may be eliminated from the DNA record stored in a comprehensive database. There are plenty of DNA markers available without getting into medical risk levels.

COLLECTING DNA IS EASY

When my son-in-law and I were teaching classes at Brandeis Summer Odyssey several years ago, he wanted his students to do a DNA project. The administrators would not allow him to take samples from students, who were minors of high school age, so they took samples from faculty members, including me. All I had to do was touch the inside of my cheek with a q-tip. Very easy and rapid. The students ran the sample through DNA testing equipment my son-in-law obtained from Harvard University. DNA samples could easily be taken at Motor Vehicle Departments when new driver's licenses are issued. They could also be taken at high schools as part of the driver's ed class.

Ira Glickstein

Monday, August 4, 2008

DNA "Fingerprints" in Anthrax Case

When anthrax was mailed to some media and political figures in 2001, resulting in some deaths, DNA testing fairly quickly determined that it was the "Ames strain," known to be the subject of experiments at Ft. Detrick. That made it most likely a rogue scientist at that facility was involved, and unlikely the killer used anthrax from a foreign source or it was "home-made." The anthrax was also in a form suitable for inhalation which apparently requires specialized processing, unlikely to be available to someone not associated with a major lab.

According to information released within the past week, some new, more detailed DNA analysis techniques have become available and affordable over the past year, and these have been used to "fingerprint" the DNA that caused the deaths and show they match the DNA used in Bruce Ivin's lab at Ft. Detrick.

While the details have not yet been released, I suspect the analysis looked for random mutations in the "junk" part of the anthrax DNA. Even assuming all the anthrax labs at Ft. Detrick started with the same exact Ames strain of anthrax, if each lab reproduced the anthrax independently, they would each end up with slight random differences that did not affect the function of the anthrax.

This case is applicable to our recent discussion of DNA "memory." You might say the DNA samples from the victims "remembered" the random mutations that occurred to their ancestors in Ivin's lab.

If you do not accept my use of the word "memory" in the Ivin's case, let me posit an analogous situation. Say a parrot has been stolen and is later recovered. The parrot is OK but, during its captivity it has added a few words to its vocabulary, and they are Hungarian words! The evidence of the parrot's memory of the Hungarian words would point to someone who speaks Hungarian as a likely suspect!

While discussing the Ivin's case with my wife, Vi, yesterday, she suggested that highly secure bio labs should consider purposely introducing telltale mutations in the junk DNA of the bioweapons in each lab. That way, if any of it was lost or stolen and used in a crime, the specific lab would be known and that could aid the investigation.

I believe there is a federal requirement that each batch of dynamite have a different mix of telltale materials added to it that are recoverable even after the dynamite has exploded. Anyone who makes a purchase must be identified and associated with the batch number of the dynamite. That way, when there is a criminal explosion involving dynamite, authorities can rapidly get a list of suspects - all those who have purchased sticks from that lot and others who are asociated with them or who may have stolen the dynamite from the legal purchaser.
Ira Glickstein

Monday, June 23, 2008

Biosemiotics


Ira suggested that I try to summarize the field of BIOSEMIOTICS ― the study of how symbol systems control living organisms and societies. I’ll try to do this in a series of short posts of less than 750 words. Then you can ask questions if I am unclear, make comments or disagree with what I have said. Hopefully, we can clear up the problems, and go on to the next post.

1ST TOPIC: WHAT IS MEMORY

A symbol system, like the genetic code, a natural language, mathematics, or an artificial computer language, requires a set of symbols and rules that reside in a memory. Memory-stored symbols are the fundamental and essential requirement for life. It is required for self-replication and all open-ended evolution. Memory is also necessary for any learning and thinking process in nervous systems. Memory is also a requirement for universal computation.

Memory and symbols can be physically implemented in endless ways, in molecules like DNA, in texts like this page, in photographs, in digital magnetic, electric and optical patterns in computers, and in neural patterns in the brain. But these particular types of memory are not what make memory of fundamental importance. So, what are the properties of memory that are essential in evolution, learning, and computation?

Two essential properties of memory are PERMANENCE and CHANGEABILITY. These properties sound incompatible, but they are complementary. In a previous post on the C- and L-minds, I compared permanence to the CONSERVATIVE aspect of memory, and I compared changeability to the LIBERAL aspect of memory. Clearly, success in adaptive evolution, learning, and social systems requires the proper balance of conservative permanence and liberal change. That is why I disagree with any liberal or conservative who claims an ideological superiority.

A good memory must also be quickly accessible, and its symbols must have the ability to effect or control a specific change. A gene must be capable of controlling protein synthesis. A brain must be capable of controlling muscles. A computer memory must be capable of changing the state of the hardware. In all of these symbol systems, genes, brains and computers, the memories have also evolved the property of self-reference. That is, genes can control their own expression, brains can think about their own thoughts (e.g., consciousness), and computer programs can address themselves. This turns out to be a mixed blessing. On the one hand, it allows organisms, brains and computers to inspect internal predictive models of the world. On the other hand, self-reference can lead to contradiction, infinite regress, and undecidable questions like whether we have free will.

While it is clear that evolution, learning, and computation could not occur unless memory has some degree of permanence and some degree of change, the nature and results of the changes are different in all three cases. In evolution, memory change is called mutation or variation, and changes are largely random. Natural selection determines the ultimate results. In nervous systems memory change is called learning. Learning is more complex and includes instruction, experience, reorganizing existing memory (thought or reasoning) and random or directed search, and often cultural selection. In computation, memory change is often called recursion or rewriting. A memory-stored program usually determines change, but programs can simulate random change and, model evolution, learning, and thinking.

Here is the classical problem of symbolic memory. The peculiar fact is that the physics of memory ― that is, the laws governing the material structures of memory symbols ― has no necessary relation to the function or meaning of the symbols. Symbol vehicles obey physical laws, but analysis of these diverse physical structures does not tell us what is important, namely the function or meaning of the symbols. Neither does analysis of these physical embodiments of memory tell us how the behaviors of memories differ in evolving organisms, brains and computers. Physical laws alone cannot predict or usefully describe the course of evolution, learning, thinking, or computation. Briefly, the problem is that symbols are arbitrarily related to their meaning or referent. The meaning or function of symbols is determined by a code or an interpreter. Symbols do not exist alone, but are a part of a language.

This fact has been a problem since the beginning of philosophy. It is
the root of the classical body-mind problem. Today in physics it is the basis of the measurement problem ― how the irreversible process interpreted as a measurement can arise from state-determined reversible laws. Some physicists also see this as an energy-information dichotomy. In biology this is the crux of the origin of life problem, how did this symbolic control of matter begin? How did molecules become messages? I call this the symbol-matter problem.
NEXT TOPIC: WHAT IS A LANGUAGE?
Howard